People don’t talk about Gemma 4 enough.
For context:
@cerebras recently made Gemma 4 31B by
@GoogleDeepMind available in public preview, and I think that the published numbers are worth testing for agentic workflows.
For my use case, it is the perfect scout model for coding agents like Oh My Pi. Not the model that makes every hard call. The model that does the parallel discovery work needed to make decisions.
Things like:
- scanning a repo
- analyzing images
- mapping a project
- summarizing sources
- triaging risks
- doing first-pass review
- finding the files a stronger model should inspect
That matters because coding agents spend a lot of time gathering context before they do the “smart” part.
The published numbers make it worth a look:
1. Speed: ~1,800 output tokens/sec
2. Cost: $0.99/M input tokens and $1.49/M output tokens
3. Quality: Analysis Intelligence Index score of 29, near Claude Haiku’s 30 on that benchmark, while delivering much higher throughput on Cerebras.
The point is not “small models beat frontier models.”
The point is routing.
The outcome for me: faster discovery loops, better context capture from text and images that previously got missed, and more workflow speed without forcing every step onto the expensive model or degrading response and code quality.
PS: At Tenex, we spend a lot of time testing models in real workflows, not benchmark slides. If you like building with new models, arguing about what they are actually good at, and turning that into product, you should apply.
Tenex Careers Page: