Piotr's picture

Piotr

piotr-ai

·

AI & ML interests

None yet

Recent Activity

liked a model about 13 hours ago

openai/gpt-oss-20b

liked a model 1 day ago

Qwen/Qwen-Image

liked a model 17 days ago

microsoft/MediPhi-Instruct

View all activity

Organizations

None yet

upvoted 2 collections 3 months ago

LLaMA-Omni

13 items • Updated May 17 • 16

NextCoder

NextCoder family of code-editing LMs developed with Selective Knowledge Transfer and its training data. • 6 items • Updated 28 days ago • 69

upvoted a paper 4 months ago

SmolVLM: Redefining small and efficient multimodal models

Paper • 2504.05299 • Published Apr 7 • 197

upvoted an article 5 months ago

Article

Introducing EuroBERT: A High-Performance Multilingual Encoder Model

By

and 3 others •

Mar 10

• 146

upvoted 2 collections 6 months ago

🧠 Reasoning datasets

Datasets with reasoning traces for math and code released by the community • 24 items • Updated May 19 • 162

SYNTHETIC-1

A collection of tasks & verifiers for reasoning datasets • 9 items • Updated 22 days ago • 62

upvoted a paper 6 months ago

Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate

Paper • 2501.17703 • Published Jan 29 • 59

upvoted a collection 7 months ago

Cosmos

The collection of Cosmos models • 31 items • Updated 16 days ago • 294

upvoted a paper 7 months ago

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining

Paper • 2501.00958 • Published Jan 1 • 107

upvoted a paper 8 months ago

Transformers Can Navigate Mazes With Multi-Step Prediction

Paper • 2412.05117 • Published Dec 6, 2024 • 5

upvoted a collection 8 months ago

Common Models

The first generation of models pretrained on Common Corpus. • 5 items • Updated Dec 5, 2024 • 39

upvoted 2 collections 9 months ago

SmolLM2

State-of-the-art compact LLMs for on-device applications: 1.7B, 360M, 135M • 16 items • Updated May 5 • 280

LayerSkip

Models continually pretrained using LayerSkip - https://arxiv.org/abs/2404.16710 • 8 items • Updated Nov 21, 2024 • 48

upvoted a collection 10 months ago

Granite 3.0 Language Models

A series of language models trained by IBM licensed under Apache 2.0 license. We release both the base pretrained and instruct models. • 8 items • Updated May 2 • 98

upvoted a paper 10 months ago

The AdEMAMix Optimizer: Better, Faster, Older

Paper • 2409.03137 • Published Sep 5, 2024 • 5

upvoted 2 collections 10 months ago

Llama 3.2

This collection hosts the transformers and original repos of the Llama 3.2 and Llama Guard 3 • 15 items • Updated Dec 6, 2024 • 628

Molmo

Artifacts for open multimodal language models. • 5 items • Updated Apr 30 • 306

upvoted 3 collections 11 months ago

Moshi v0.1 Release

MLX, Candle & PyTorch model checkpoints released as part of the Moshi release from Kyutai. Run inference via: https://github.com/kyutai-labs/moshi • 15 items • Updated Apr 18 • 235

Qwen2.5

Qwen2.5 language models, including pretrained and instruction-tuned models of 7 sizes, including 0.5B, 1.5B, 3B, 7B, 14B, 32B, and 72B. • 46 items • Updated 16 days ago • 630

DataGemma Release

A series of pioneering open models that help ground LLMs in real-world data through Data Commons. • 2 items • Updated 27 days ago • 87