Build agentic workflows completely offline.
The Antigravity SDK now supports local execution with Gemma 4 and LiteRT. Run agents entirely on your local machine with:
๐ต Zero token costs
๐ Total data privacy
๐ Offline reliability
Bonus feature: Support for OpenAI-compatible endpoints. Use Ollama, llama.cpp, vLLM and more to serve Gemma ๐ช
Get started: pip install google-antigravity litert-lm
Read the details:
Show more
"DiffusionGemma as Jev" showcases the power of non-autoregressive architectures.
While Jev demonstrates the value of rapid decision models, running DiffusionGemma in this paradigm leverages canvas diffusion to evaluate structured choices in a single parallel pass:
โก ๏ธMassive Parallelism: Denoises across an open canvas in a single step instead of sequential autoregressive token generation (~0.2s on a DGX spark).
๐ง Full Bidirectional Attention: Allows every option to attend to the full context concurrently, yielding well-calibrated decision distributions.
๐๏ธ Multimodal Grounding: Inherits Gemma 4's spatial vision capabilities for complex visual and text decisions.
Read more about this approach here:
Show more
I ran some real, live evals on Jev vs DiffusionGemma-as-Jev (my patch for vLLM!)
DiffusionGemma comes out as the winner, I think.
Headlines:
Is Jev faster than DiffusionGemma? No โ (API vs DGX Spark)
Is Jev smarter than DiffusionGemma? No โ (they're roughly tied!)
Show more
We prototyped a tactile desk robot using LEGO, a Raspberry Pi, and a hybrid LLM setup:
โจ Gemma 4 runs locally for instant, private zero-cost chat
โจ Auto-escalates to Gemini Flash for complex reasoning & code
Hereโs how we wired and coded DinoDesk AI:
Show more
Standup Pulse: an open-source project running async Slack standups using Gemma 4 26B-A4B (GGUF via llama.cpp) on an Apple M5 Max. It pairs Mastra for typed tool selection with CopilotKit Channels for restrained Slack Block Kit (Slack's native layout format) interactions while keeping all model inference, standup records, and traces local in SQLite.
This architecture is a great demonstration of how to connect Slack to a project without exposing the local server!
๐ Blog:
๐ Repo:
Show more
You don't need to be a command-line expert to run local AI.
wraps the heavy lifting of llama.cpp into a clean UI, making it simple to deploy and manage models like Gemma 4. Instead of dealing with config files, you get one-click model downloads, clear memory estimates, and no coding required!
Watch the video to see Gemma 4 in action as it parses tables from receipts, streams reasoning logs, and connects to MCP for web search.
Show more
ICYMI, we're celebrating 1 BILLION+ downloads for
@googlegemma ๐
How are developers actually using open models?
@GoogleDeepMindโs
@DynamicWebPaige caught up with devs and collaborators like
@UnslothAI and
@Qualcomm to hear how theyโre building on-device tools, running local fine-tuning, and pushing multimodal breakthroughs.
Show more
Gemma is only as powerful as the community building with it.
We launched awesome-gemma to highlight amazing community-made tools, research, and applications in the ecosystem.
What have you built?
Drop a link below or submit a PR to get featured:
Show more
You don't always need a frontier model.
A recent benchmark found that Gemma 4 31B matches Sonnet 5 on answer quality at ~40x lower cost.
With high cost-efficiency and low latency, Gemma unlocks high-volume use cases that are uneconomical with larger models.
Show more
๐ 1 BILLION DOWNLOADS ๐
To celebrate this exciting milestone, weโre hosting an exclusive evening in SF on Aug 20 dedicated to YOU, the open-source builders, researchers, and contributors driving the Gemmaverse forward.
Space is limited. Apply for your spot here:
Show more
This week, Gemma surpassed 900 million downloads! ๐
All the way from Gemma 1 and ShieldGemma to MedGemma and Gemma 4, we'll keep supporting open source. More to come!
Gemma 4 just crossed 300 million downloads.
Thank you to the developers, researchers, and open-source community building with us. Your work and feedback drive this project forward.
Let's keep building!
Show more
Extremely fast multimodal inference!
Damage Scout uses Gemma 4 on
@cerebras running at an impressive 2,300+ toks/s! It analyzes rental car walkaround videos and generates annotated damage reports with box coordinates in <6 seconds.
Show more
Voice AI without the wait! โฑ๏ธ
Thanks to Hugging Face and Cerebras, developers can now use the Gemma 4 31B model as the brain for voice AI at ultra-fast inference speeds. Add it to a fully open-source, cascaded speech-to-speech stack that can be used to power existing voice apps! ๐ฃ๏ธ
Show more
Weโre rolling out some big improvements to Gemma 4, fueled by incredible community feedback and contributions!
Here is a breakdown of whatโs being fixed and updated in this release: ๐งต๐
Show more
Gemma 4 is live on
@cerebras, the fastest multimodal inference ever!
Running on Gemma 4 31B open-weight model at a blistering 1,500+ tokens/sec. That's a 15x speedup, unlocking real-time visual and agentic loops without the GPU lag.
Show more
Hugging Face Gemma Challenge results are in! ๐
Over 6 days, more than 100 AI agents and humans collaborated to make Gemma 4 inference 5x faster on a single NVIDIA A10G GPU.
- Fastest result: 491.8 TPS (fastest overall, but resulted in a drop in model quality in other areas)
- Fastest lossless: 315 TPS
A great example of what humans and agents can achieve when they work together.
Show more
Gemma 4 31B at over 1,800 tokens per second!
Gemma 4 is now in Public Preview on Cerebras.
Gemma 4 is the first multimodal model on Cerebras! ๏ธ
What can you build with Gemma 4 31B running at 1500 tokens per second?
Join the Cerebras x Gemma 4 24-hour virtual hackathon this Sunday to compete for $5,000 in prizes.
Participants get early access to Gemma 4 on Cerebras.
Show more
Gemma 4 just hit 200M downloads in only 2.5 months!
For context, total downloads across the entire Gemma family of models were at 100M when we launched Gemma 3. The community's acceleration is incredible. Thank you to everyone building with Gemma.
Watch how developers are driving real-world impact:
Show more
16 parallel runs of Gemma 4 26B A4B on a single NVIDIA DGX Spark!
Pushing 18 tok/s per instance and a 300 tok/s aggregate. It can even hit 32 parallel runs.
This level of concurrency highlights how efficient the architecture is.
Show more