Register and share your invite link to earn from video plays and referrals.

Google Gemma
@googlegemma
The official home of Google's Gemma. Lightweight, state-of-the-art open models by Google DeepMind, built on Gemini tech. What will you build? ๐Ÿš€๐Ÿ’ป
0 Following    101.1K Followers
Build agentic workflows completely offline. The Antigravity SDK now supports local execution with Gemma 4 and LiteRT. Run agents entirely on your local machine with: ๐Ÿ’ต Zero token costs ๐Ÿ”’ Total data privacy ๐Ÿ”Œ Offline reliability Bonus feature: Support for OpenAI-compatible endpoints. Use Ollama, llama.cpp, vLLM and more to serve Gemma ๐Ÿ’ช Get started: pip install google-antigravity litert-lm Read the details:
Show more
"DiffusionGemma as Jev" showcases the power of non-autoregressive architectures. While Jev demonstrates the value of rapid decision models, running DiffusionGemma in this paradigm leverages canvas diffusion to evaluate structured choices in a single parallel pass: โšก ๏ธMassive Parallelism: Denoises across an open canvas in a single step instead of sequential autoregressive token generation (~0.2s on a DGX spark). ๐Ÿง  Full Bidirectional Attention: Allows every option to attend to the full context concurrently, yielding well-calibrated decision distributions. ๐Ÿ‘๏ธ Multimodal Grounding: Inherits Gemma 4's spatial vision capabilities for complex visual and text decisions. Read more about this approach here:
Show more
I ran some real, live evals on Jev vs DiffusionGemma-as-Jev (my patch for vLLM!) DiffusionGemma comes out as the winner, I think. Headlines: Is Jev faster than DiffusionGemma? No โŒ (API vs DGX Spark) Is Jev smarter than DiffusionGemma? No โŒ (they're roughly tied!)
Show more
0
57
4.4K
524
Forward to community
We prototyped a tactile desk robot using LEGO, a Raspberry Pi, and a hybrid LLM setup: โœจ Gemma 4 runs locally for instant, private zero-cost chat โœจ Auto-escalates to Gemini Flash for complex reasoning & code Hereโ€™s how we wired and coded DinoDesk AI:
Show more
Standup Pulse: an open-source project running async Slack standups using Gemma 4 26B-A4B (GGUF via llama.cpp) on an Apple M5 Max. It pairs Mastra for typed tool selection with CopilotKit Channels for restrained Slack Block Kit (Slack's native layout format) interactions while keeping all model inference, standup records, and traces local in SQLite. This architecture is a great demonstration of how to connect Slack to a project without exposing the local server! ๐Ÿ”— Blog: ๐Ÿ”— Repo:
Show more
You don't need to be a command-line expert to run local AI. wraps the heavy lifting of llama.cpp into a clean UI, making it simple to deploy and manage models like Gemma 4. Instead of dealing with config files, you get one-click model downloads, clear memory estimates, and no coding required! Watch the video to see Gemma 4 in action as it parses tables from receipts, streams reasoning logs, and connects to MCP for web search.
Show more
ICYMI, we're celebrating 1 BILLION+ downloads for @googlegemma ๐Ÿ’Ž How are developers actually using open models? @GoogleDeepMindโ€™s @DynamicWebPaige caught up with devs and collaborators like @UnslothAI and @Qualcomm to hear how theyโ€™re building on-device tools, running local fine-tuning, and pushing multimodal breakthroughs.
Show more
Gemma is only as powerful as the community building with it. We launched awesome-gemma to highlight amazing community-made tools, research, and applications in the ecosystem. What have you built? Drop a link below or submit a PR to get featured:
Show more
You don't always need a frontier model. A recent benchmark found that Gemma 4 31B matches Sonnet 5 on answer quality at ~40x lower cost. With high cost-efficiency and low latency, Gemma unlocks high-volume use cases that are uneconomical with larger models.
Show more
0
178
2.8K
147
Forward to community
๐ŸŽ‰ 1 BILLION DOWNLOADS ๐ŸŽ‰ To celebrate this exciting milestone, weโ€™re hosting an exclusive evening in SF on Aug 20 dedicated to YOU, the open-source builders, researchers, and contributors driving the Gemmaverse forward. Space is limited. Apply for your spot here:
Show more
This week, Gemma surpassed 900 million downloads! ๐ŸŽ‰ All the way from Gemma 1 and ShieldGemma to MedGemma and Gemma 4, we'll keep supporting open source. More to come!
0
73
1.7K
126
Forward to community
Gemma 4 just crossed 300 million downloads. Thank you to the developers, researchers, and open-source community building with us. Your work and feedback drive this project forward. Let's keep building!
Show more
0
64
1.7K
123
Forward to community
Extremely fast multimodal inference! Damage Scout uses Gemma 4 on @cerebras running at an impressive 2,300+ toks/s! It analyzes rental car walkaround videos and generates annotated damage reports with box coordinates in <6 seconds.
Show more
Voice AI without the wait! โฑ๏ธ Thanks to Hugging Face and Cerebras, developers can now use the Gemma 4 31B model as the brain for voice AI at ultra-fast inference speeds. Add it to a fully open-source, cascaded speech-to-speech stack that can be used to power existing voice apps! ๐Ÿ—ฃ๏ธ
Show more
0
52
2.5K
246
Forward to community
Weโ€™re rolling out some big improvements to Gemma 4, fueled by incredible community feedback and contributions! Here is a breakdown of whatโ€™s being fixed and updated in this release: ๐Ÿงต๐Ÿ‘‡
Show more
0
131
3.3K
325
Forward to community
Gemma 4 is live on @cerebras, the fastest multimodal inference ever! Running on Gemma 4 31B open-weight model at a blistering 1,500+ tokens/sec. That's a 15x speedup, unlocking real-time visual and agentic loops without the GPU lag.
Show more
0
51
1.3K
104
Forward to community
Hugging Face Gemma Challenge results are in! ๐Ÿ“ˆ Over 6 days, more than 100 AI agents and humans collaborated to make Gemma 4 inference 5x faster on a single NVIDIA A10G GPU. - Fastest result: 491.8 TPS (fastest overall, but resulted in a drop in model quality in other areas) - Fastest lossless: 315 TPS A great example of what humans and agents can achieve when they work together.
Show more
0
62
1.7K
162
Forward to community
Gemma 4 31B at over 1,800 tokens per second! Gemma 4 is now in Public Preview on Cerebras.
0
39
1.7K
96
Forward to community
Gemma 4 is the first multimodal model on Cerebras! ๏ธ What can you build with Gemma 4 31B running at 1500 tokens per second? Join the Cerebras x Gemma 4 24-hour virtual hackathon this Sunday to compete for $5,000 in prizes. Participants get early access to Gemma 4 on Cerebras.
Show more
0
46
1.1K
106
Forward to community
Gemma 4 just hit 200M downloads in only 2.5 months! For context, total downloads across the entire Gemma family of models were at 100M when we launched Gemma 3. The community's acceleration is incredible. Thank you to everyone building with Gemma. Watch how developers are driving real-world impact:
Show more
0
113
1.7K
178
Forward to community
16 parallel runs of Gemma 4 26B A4B on a single NVIDIA DGX Spark! Pushing 18 tok/s per instance and a 300 tok/s aggregate. It can even hit 32 parallel runs. This level of concurrency highlights how efficient the architecture is.
Show more
0
89
2.6K
232
Forward to community