Gemma 4 just hit 200M downloads in only 2.5 months!
For context, total downloads across the entire Gemma family of models were at 100M when we launched Gemma 3. The community's acceleration is incredible. Thank you to everyone building with Gemma.
Watch how developers are driving real-world impact:
Gemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️
MTP enables Google Gemma 4 run ~1.4–2.2× faster with no accuracy loss.
Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s.
GGUFs + Guide:
Gemma 4 models are sized to run efficiently & can be fine-tuned for state-of-the-art performance on specific tasks. We’ve already seen success with this approach, including our research with Yale University to discover new pathways for cancer therapy.
Gemma 4: Now up to 3x Faster. ⚡
Same quality, way more speed. Our new MTP drafters allow Gemma 4 to predict multiple tokens at once, effectively tripling your output speed without compromising intelligence.