Google dropped MTP versions of Gemma4. Ran them on my DGX Spark.
The 31B dense model went from 3.94 → 8.91 tok/s. That's +126%.
Full results:
[26B A4B]
> 25.24 → 31.69 tok/s (+25.6%)
> TTFT 755 → 332ms (-56%)
[31B]
> 3.94 → 8.91 tok/s (+126%)
> TTFT 599 → 378ms (-37%)
If you're not running MTP, you're leaving free perf on the table.
We're working on making the local model experience better in Hermes, what are the best local models at each weight class?
My blind guess, please correct:
8-16 GB VRAM
Gemma4 12B
24-32 GB VRAM
Qwen3.6 27B
Qwen3.6 35B
128 GB VRAM (Spark, M3 Max)
??? Can you do DSv4-Flash?
🚀 ZDTaichu5.0-9B is now on ModelScope! 🤖
An on-device multimodal model from TaichuAI. At 9B parameters it runs on a single GPU and brings spatial reasoning, embodied AI and agentic tool use to edge deployment. Qwen3.5-9B backbone + C-RADIOv4-H vision encoder, 128K context, any-resolution image and video input.
🧭 Spatial reasoning: leads the compared 10B-scale open VLMs (Qwen3.5-9B, STEP3-VL-10B, gemma4-8B-E4B) and scores above Gemini 3 Pro, Grok 4 and GPT-5.2 on ViewSpatial, MMSI-Bench and MindCube-tiny
🛠️ Agent: highest among the compared open models on TAU2-Bench, Claw-Eval and IFEval
📄 First-tier results on documents, charts, OCR, visual math and video, with a ready-to-use vLLM branch and Docker image
🧠 Entropy-Gated Adaptive Recurrent Reasoning: extra latent refinement steps go only to the hard tokens
Gemma is only as powerful as the community building with it.
We launched awesome-gemma to highlight amazing community-made tools, research, and applications in the ecosystem.
What have you built?
Drop a link below or submit a PR to get featured:
Gemma 4 Mobile is now available on iPhone and iPad!
This model is a variant optimized for mobile use cases and comes with a reduced memory footprint. Running natively with MLX.
And congrats to @googlegemma on reaching 1 billion downloads! 🎉
Gemma 4 just hit 200M downloads in only 2.5 months!
For context, total downloads across the entire Gemma family of models were at 100M when we launched Gemma 3. The community's acceleration is incredible. Thank you to everyone building with Gemma.
Watch how developers are driving real-world impact: