DeepSeek-V4-Flash can now run 2× faster locally with DSpark! ⚡️
DSpark enables V4-Flash-0731 GGUFs to generate ~1.4–2× faster with no accuracy change.
DeepSeek-V4-Flash-0731 can reach at 120 tokens/s.
GGUFs:
Guide:
Show more
Qwen3.8-27B is coming! 🔥
Will run locally on 17GB RAM/VRAM setups.
DeepSeek V4 Flash 0731 can now be run locally! 🐳
Run DeepSeek V4 Flash lossless 4-bit on 168GB RAM and 3-bit on 110GB RAM.
V4 Flash 0731 outperforms V4 Pro. Run via Unsloth or llama.cpp. Smaller quants coming today.
Guide:
GGUF:
Show more
You can now run Inkling-Small, a new 276B model by Thinking Machines.
Inkling-Small is the strongest open model for its size and runs local on 128GB RAM.
Apache-2.0 Licensed, it has image, audio + 1M context support.
Guide:
GGUF:
Show more
Kimi K3 can now be run locally! ✨
The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% size).
Run on a Mac Studio + 128GB RAM device.
Kimi K3 is the strongest open model to date.
Guide:
GGUF:
Show more
Introducing Unsloth for AMD 🚀
You can now train & run LLMs on your AMD hardware
• We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs
• Works on Windows, WSL, Linux
• Train Qwen, Gemma on 3GB VRAM
GitHub:
Works on Radeon, Instinct, Ryzen and data center GPUs with up to 2× faster with 70% less VRAM and no accuracy loss via our custom Triton kernels and math algorithms. We also support optimized ROCm builds for GGUF & Safetensors inference.
Unsloth is an open-source local UI for faster LLM training and inference, with tool-call healing, code execution, secure web search, remote APIs, and HTTPS deployment. Connect local models to Claude Code, Codex agents and run the latest Kimi, GLM, DeepSeek, Qwen3.6, and Gemma 4 models.
🔗Blog + Guide:
Show more
You can now run Thinking Machines Inkling!
Inkling is a 975B open model with image, audio and 1M context support.
We quantized Inkling to Dynamic 1-bit (-86% size) and retained 74.2% of top-1% accuracy. Run on 280GB.
Guide:
GGUF:
Show more
We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU.
Qwen3.6-27B NVFP4 runs on 24GB VRAM.
35B-A3B can hit 17,561 tok/s (B200).
We also improved accuracy, tool calling, agent use, and looping.
Guide:
Qwen3.6 NVFP4:
Show more
What’s your go-to local model right now?
1-bit GLM-5.2 GGUF vs. Claude 4.8 Opus vs. GPT-5.5
We gave 3 models the same prompt and compared one-shot outputs.
The 1-bit GLM-5.2 GGUF ran locally on a Mac Studio M3 Ultra with 256GB RAM at ~21.6 tok/s.
Which output do you like best?
GGUF:
Show more
GLM-5.2 can now be run locally!🔥
The 2-bit model retains ~82% accuracy after we shrunk it from 1.51TB to 238GB (-84% size).
Run on a 256GB Mac or RAM/VRAM setups.
GLM-5.2 is the strongest open model to date.
Guide:
GGUF:
Show more
GLM-5.2 can now be run locally!🔥
The 2-bit model retains ~82% accuracy after we shrunk it from 1.51TB to 238GB (-84% size).
Run on a 256GB Mac or RAM/VRAM setups.
GLM-5.2 is the strongest open model to date.
Guide:
GGUF:
Show more
You can now run Kimi K2.7 Code locally! 🌘
We shrank the 1T model to 325GB (-48%) via Dynamic 2-bit where important layers are upcasted.
Run at >40 tok/s on 330GB RAM/VRAM setups.
Run full precision on 610 GB.
Guide:
GGUF:
Show more
🌘 Kimi-K2.7-Code, our latest coding model, is now released and open-sourced!
🔷 Improved coding & agent performance over K2.6: +21.8% on Kimi Code Bench v2, +11.0% on Program Bench, and +31.5% on MLS Bench Lite.
🔷 Reasoning efficiency: Less overthinking, with 30% lower reasoning-token usage compared to K2.6.
🔷 Long-horizon coding: Improved instruction following, higher end-to-end coding task success rates.
⚡️ 6x High-Speed Mode coming soon!
🔌 Available today via Kimi API and Kimi Code.
🔗 Kimi Code:
🔗 API:
Show more
Local AI in action! MiniMax M3 unning locally on a single M3 Ultra 512GB in Unsloth Studio! 🔥
Here UD-Q5_K_XL decoding at 32.5 toks/s!
MiniMax M3 can now be run locally!🔥
MiniMax-M3 is a new 428B (23B active) open model with 1M context that performs on par with Gemini 3.1 Pro.
Run Dynamic 2-bit GGUF on 138GB RAM/VRAM or 3-bit on 165GB.
GGUF:
Guide:
Show more
MiniMax M3, Open-Weight, Now On Hugging Face , with only ~428B parameters and ~23B activated parameters
Weights:
MiniMax Sparse Attention:
DiffusionGemma can now run at 2000+ tokens/sec! ⚡
We made local DiffusionGemma inference 1.8× faster.
Run it on 18GB RAM via Unsloth Studio.
GitHub:
Guide:
Show more
Google releases DiffusionGemma.✨
The new 26B-A4B diffusion text model runs locally on 18GB RAM.
It supports high-speed text generation, thinking, image, video and 256K context.
Run and train via Unsloth Studio.
GGUF:
Guide:
Show more
Gemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️
MTP enables Google Gemma 4 run ~1.4–2.2× faster with no accuracy loss.
Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s.
GGUFs + Guide:
Show more
Google releases DiffusionGemma.✨
The new 26B-A4B diffusion text model runs locally on 18GB RAM.
It supports high-speed text generation, thinking, image, video and 256K context.
Run and train via Unsloth Studio.
GGUF:
Guide:
Show more
Meet DiffusionGemma!
An experimental open model that explores a fast approach to text generation, released under an Apache 2.0 license.
Moving beyond sequential, token-by-token processes to generate entire blocks of text simultaneously. Here’s what’s new with DiffusionGemma: 👇
Show more
Google releases Gemma 4 QAT. ✨
You can now run Gemma 4 at 3x less memory with near original performance.
Quantization-Aware Training (QAT) makes it possible to run Gemma 4 26B-A4B on 16GB RAM.
GGUFs:
QAT Guide:
Show more
We just dropped Gemma 4 Quantization-Aware Training (QAT) checkpoints on Hugging Face!
All Gemma 4 model sizes and their drafters are now optimized with QAT to cut memory requirements and maximize on-device performance!
Show more
You can now run NVIDIA Nemotron 3 Ultra, a new 550B open model.
Nemotron-3-Ultra-550B-A55B is NVIDIA's largest LLM yet, with 1M context, frontier coding & chat.
Run 2-bit on 200GB RAM, 3-bit on 256GB, 8-bit on 600GB.
GGUF:
Guide:
Show more
Today we're shipping Nemotron 3 Ultra.
A 550B MoE frontier-intelligence open model built for long-running agents.
It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models.
Show more
2-bit Gemma 4 12B GGUF, only 4.66 GB on disk, managed to cite 15 sites from a single prompt.
Try this locally on >6GB RAM via Unsloth Studio.
GitHub:
Show more
Gemma 4 12B can now run locally on just 8GB RAM via Dynamic GGUFs.
Google's new model, Gemma 4 12B Unified supports image, audio and 256K context.
You can run and train the model via Unsloth Studio.
GGUF:
Guide:
Show more