Register and share your invite link to earn from video plays and referrals.

Unsloth AI
@UnslothAI
Train and run models locally! 🦥
479 Following    87.2K Followers
DeepSeek-V4-Flash can now run 2× faster locally with DSpark! ⚡️ DSpark enables V4-Flash-0731 GGUFs to generate ~1.4–2× faster with no accuracy change. DeepSeek-V4-Flash-0731 can reach at 120 tokens/s. GGUFs: Guide:
Show more
0
85
2.5K
221
Forward to community
Qwen3.8-27B is coming! 🔥 Will run locally on 17GB RAM/VRAM setups.
0
242
7K
526
Forward to community
DeepSeek V4 Flash 0731 can now be run locally! 🐳 Run DeepSeek V4 Flash lossless 4-bit on 168GB RAM and 3-bit on 110GB RAM. V4 Flash 0731 outperforms V4 Pro. Run via Unsloth or llama.cpp. Smaller quants coming today. Guide: GGUF:
Show more
0
120
2.5K
332
Forward to community
You can now run Inkling-Small, a new 276B model by Thinking Machines. Inkling-Small is the strongest open model for its size and runs local on 128GB RAM. Apache-2.0 Licensed, it has image, audio + 1M context support. Guide: GGUF:
Show more
Kimi K3 can now be run locally! ✨ The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% size). Run on a Mac Studio + 128GB RAM device. Kimi K3 is the strongest open model to date. Guide: GGUF:
Show more
0
433
10.5K
1.4K
Forward to community
Introducing Unsloth for AMD 🚀 You can now train & run LLMs on your AMD hardware • We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs • Works on Windows, WSL, Linux • Train Qwen, Gemma on 3GB VRAM GitHub: Works on Radeon, Instinct, Ryzen and data center GPUs with up to 2× faster with 70% less VRAM and no accuracy loss via our custom Triton kernels and math algorithms. We also support optimized ROCm builds for GGUF & Safetensors inference. Unsloth is an open-source local UI for faster LLM training and inference, with tool-call healing, code execution, secure web search, remote APIs, and HTTPS deployment. Connect local models to Claude Code, Codex agents and run the latest Kimi, GLM, DeepSeek, Qwen3.6, and Gemma 4 models. 🔗Blog + Guide:
Show more
0
68
1.4K
178
Forward to community
You can now run Thinking Machines Inkling! Inkling is a 975B open model with image, audio and 1M context support. We quantized Inkling to Dynamic 1-bit (-86% size) and retained 74.2% of top-1% accuracy. Run on 280GB. Guide: GGUF:
Show more
We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. Qwen3.6-27B NVFP4 runs on 24GB VRAM. 35B-A3B can hit 17,561 tok/s (B200). We also improved accuracy, tool calling, agent use, and looping. Guide: Qwen3.6 NVFP4:
Show more
0
169
2.8K
348
Forward to community
What’s your go-to local model right now?
0
322
582
19
Forward to community
1-bit GLM-5.2 GGUF vs. Claude 4.8 Opus vs. GPT-5.5 We gave 3 models the same prompt and compared one-shot outputs. The 1-bit GLM-5.2 GGUF ran locally on a Mac Studio M3 Ultra with 256GB RAM at ~21.6 tok/s. Which output do you like best? GGUF:
Show more
GLM-5.2 can now be run locally!🔥 The 2-bit model retains ~82% accuracy after we shrunk it from 1.51TB to 238GB (-84% size). Run on a 256GB Mac or RAM/VRAM setups. GLM-5.2 is the strongest open model to date. Guide: GGUF:
Show more
0
172
3.6K
411
Forward to community
GLM-5.2 can now be run locally!🔥 The 2-bit model retains ~82% accuracy after we shrunk it from 1.51TB to 238GB (-84% size). Run on a 256GB Mac or RAM/VRAM setups. GLM-5.2 is the strongest open model to date. Guide: GGUF:
Show more
0
274
7.4K
896
Forward to community
You can now run Kimi K2.7 Code locally! 🌘 We shrank the 1T model to 325GB (-48%) via Dynamic 2-bit where important layers are upcasted. Run at >40 tok/s on 330GB RAM/VRAM setups. Run full precision on 610 GB. Guide: GGUF:
Show more
🌘 Kimi-K2.7-Code, our latest coding model, is now released and open-sourced! 🔷 Improved coding & agent performance over K2.6: +21.8% on Kimi Code Bench v2, +11.0% on Program Bench, and +31.5% on MLS Bench Lite. 🔷 Reasoning efficiency: Less overthinking, with 30% lower reasoning-token usage compared to K2.6. 🔷 Long-horizon coding: Improved instruction following, higher end-to-end coding task success rates. ⚡️ 6x High-Speed Mode coming soon! 🔌 Available today via Kimi API and Kimi Code. 🔗 Kimi Code: 🔗 API:
Show more
0
172
2.9K
304
Forward to community
Local AI in action! MiniMax M3 unning locally on a single M3 Ultra 512GB in Unsloth Studio! 🔥 Here UD-Q5_K_XL decoding at 32.5 toks/s!
MiniMax M3 can now be run locally!🔥 MiniMax-M3 is a new 428B (23B active) open model with 1M context that performs on par with Gemini 3.1 Pro. Run Dynamic 2-bit GGUF on 138GB RAM/VRAM or 3-bit on 165GB. GGUF: Guide:
Show more
MiniMax M3, Open-Weight, Now On Hugging Face , with only ~428B parameters and ~23B activated parameters Weights: MiniMax Sparse Attention:
DiffusionGemma can now run at 2000+ tokens/sec! ⚡ We made local DiffusionGemma inference 1.8× faster. Run it on 18GB RAM via Unsloth Studio. GitHub: Guide:
Show more
Google releases DiffusionGemma.✨ The new 26B-A4B diffusion text model runs locally on 18GB RAM. It supports high-speed text generation, thinking, image, video and 256K context. Run and train via Unsloth Studio. GGUF: Guide:
Show more
0
64
1.7K
186
Forward to community
Gemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️ MTP enables Google Gemma 4 run ~1.4–2.2× faster with no accuracy loss. Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s. GGUFs + Guide:
Show more
0
62
2.2K
259
Forward to community
Google releases DiffusionGemma.✨ The new 26B-A4B diffusion text model runs locally on 18GB RAM. It supports high-speed text generation, thinking, image, video and 256K context. Run and train via Unsloth Studio. GGUF: Guide:
Show more
Meet DiffusionGemma! An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license. Moving beyond sequential, token-by-token processes to generate entire blocks of text simultaneously. Here’s what’s new with DiffusionGemma: 👇
Show more
0
66
1.8K
250
Forward to community
Google releases Gemma 4 QAT. ✨ You can now run Gemma 4 at 3x less memory with near original performance. Quantization-Aware Training (QAT) makes it possible to run Gemma 4 26B-A4B on 16GB RAM. GGUFs: QAT Guide:
Show more
We just dropped Gemma 4 Quantization-Aware Training (QAT) checkpoints on Hugging Face! All Gemma 4 model sizes and their drafters are now optimized with QAT to cut memory requirements and maximize on-device performance!
Show more
0
93
2.9K
411
Forward to community
You can now run NVIDIA Nemotron 3 Ultra, a new 550B open model. Nemotron-3-Ultra-550B-A55B is NVIDIA's largest LLM yet, with 1M context, frontier coding & chat. Run 2-bit on 200GB RAM, 3-bit on 256GB, 8-bit on 600GB. GGUF: Guide:
Show more
Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models.
Show more
2-bit Gemma 4 12B GGUF, only 4.66 GB on disk, managed to cite 15 sites from a single prompt. Try this locally on >6GB RAM via Unsloth Studio. GitHub:
Show more
Gemma 4 12B can now run locally on just 8GB RAM via Dynamic GGUFs. Google's new model, Gemma 4 12B Unified supports image, audio and 256K context. You can run and train the model via Unsloth Studio. GGUF: Guide:
Show more
0
46
1.6K
193
Forward to community