Register and share your invite link to earn from video plays and referrals.

Daniel Han
@danielhanchen
Building @UnslothAI • Making open-source LLMs faster, better & more accessible • YC S24 • ex-NVIDIA ML
2K Following    36.3K Followers
Transformers has supported loading GGUF files for a few years now, by unquantizing them. Thanks to @_marcsun, we're now using GGML kernels through the `kernels` library to run at the same performance as llama.cpp Huge kudos to the entire @ggml_org for making these kernels!
Show more
We got @UnslothAI a DGX Station! @DanielHanChen and @NaderLikeLadder checked out Unsloth’s new @Dell Pro Max with GB300 and talked about what comes next: support for more models, faster quantization, and more efficient reinforcement learning.
Show more
In 2026 Local AI adoption has dramatically grown and outpaced all our own predictions at Unsloth! We're working on a lot of cool stuff which will make local AI even faster and more accessible in the coming weeks!
Show more
Unsloth has surpassed 500M model downloads on Hugging Face! 🦥🤗 Qwen3.8-27B GGUF is already Unsloth’s #1# most-downloaded model ever. Thanks for all your support!
We just did a hotfix / mini update to Unsloth Desktop! 1. Added image editing to Qwen-Image-2.1 2. Fixed diffusers updating issues and fixed GGUFs not loading for Qwen-Image 3. Fixed diffusion black artifacts for A100 and consumer GPUs You should receive a banner for updating
Show more
Qwen-Image-2.1 FP8 and GGUF quants should now run properly in Unsloth Desktop! 💜 Image gen and editing are both supported. GitHub:
Qwen-Image-2.1 works great in Unsloth Desktop via INT8 / FP8 and GGUFs! Pinned offloading to RAM also allows INT8 / FP8 to fit in under 6-8GB of VRAM, and is still relatively fast! We also made some dynamic GGUFs for it as well!
Show more
Qwen-Image-2.1 can now run locally on 12GB VRAM with Unsloth GGUFs! 🖼️ The 7B model performs on par with Nano Banana 2.0. For higher quality, you can also run Dynamic FP8 on just 6GB of VRAM via offloading. GGUF: Guide:
Show more
We added multi user accounts to @UnslothAI Desktop, a revamped Docker & custom Jupyter Notebook with custom themes, titles and expandable cells, RDNA1+2 support, ARM64 Windows CUDA support, faster GRPO, and FP8/INT8 image diffusion support for 2x faster inference!
Show more
You can now train and run 500+ models locally with our Unsloth Docker image! 🐳 Use our new GUI or notebooks workflow. No setup required. Works on NVIDIA and AMD. Guide: GitHub:
Show more
You can now train and run 500+ models locally with our Unsloth Docker image! 🐳 Use our new GUI or notebooks workflow. No setup required. Works on NVIDIA and AMD. Guide: GitHub:
Show more
Introducing Unsloth Desktop 🦥 The first desktop app to run and train models locally. • Open-source. Runs on Mac, Windows and Linux • Supports MLX, diffusion image/video, audio, GGUF • Connect Claude Code and Codex to local LLMs • 50% more accurate, self-healing tool calls + sandboxed code exec • Works for CPU + multiGPU setups - NVIDIA, AMD, Intel, Mac • Train models 2× faster with 70% less VRAM • Private web search, deep research, RAG, MCP and exports (NVFP4, GGUF) • Use Unsloth’s OpenAI-compatible API and cloud models • Securely deploy LLMs remotely and access anywhere Unsloth Desktop is now available on and GitHub. GitHub: Blog and Guide:
Show more
@bnjmn_marie We're also working on a smaller version for NVFP4 since folks want to squeeze more context lengths out on a 5090 :)
Qwen3.8-27B NVFP4 variants are very close in accuracy. The real differences are memory and speed. If you're VRAM-limited, minima-ai/mnma_qwen3.8_27b_nvfp4 is a good pick, but it doesn't include MTP for faster inference. NVIDIA's version has MTP, but in my long-context coding tests MTP-4 is only ~2× faster than no MTP, and still ~2.5–3× slower than Unsloth (RTX Pro 6000). A likely reason: NVIDIA quantizes lm_head to NVFP4, while Unsloth keeps it FP8. Since MTP shares the target model's lm_head, this can hurt prediction quality and acceptance rate. So my pick is Unsloth NVFP4. Now testing accuracy on long-horizon agentic coding.
Show more
Qwen3.8-27B is now the most-liked GGUF ever. Soon it'll be the top 30 most-liked model overall. 🤯 It's all thanks to you all, HF, and the Qwen team!
Qwen3.8-27B Unsloth GGUF is now the #1# most-liked GGUF of all time! The model hit 10M downloads and 3.7K likes in just 24 days on Hugging Face - all thanks to you. 🤗🦥 GGUF: Guide:
Show more
Qwen3.8-27B Unsloth GGUF is now the #1# most-liked GGUF of all time! The model hit 10M downloads and 3.7K likes in just 24 days on Hugging Face - all thanks to you. 🤗🦥 GGUF: Guide:
Show more
0
73
1.7K
124
Forward to community
ICYMI, we're celebrating 1 BILLION+ downloads for @googlegemma 💎 How are developers actually using open models? @GoogleDeepMind’s @DynamicWebPaige caught up with devs and collaborators like @UnslothAI and @Qualcomm to hear how they’re building on-device tools, running local fine-tuning, and pushing multimodal breakthroughs.
Show more
Get faster inference with GLM-5.3-Flash GGUFs out of the box in Unsloth Desktop. We enabled MTP and faster long context decoding!
We made GLM-5.3-Flash run 3.3x faster locally! Local GGUF inference is now 1.6–3.4× faster with optimized decoding and bonus multi-token prediction. Run 3-bit on 128GB setups via Unsloth Desktop or llama.cpp. Guide: GGUF:
Show more
Hermes now has a local backend! There's support for Unsloth UD-Q4_K_XL and UD-Q4_K_M quants for DeepSeek-V4-Flash, Qwen3.8-27B, 3.6-35B-A3B and Qwen3.8-Flash-Next in Hermes!
You can now run Unsloth GGUFs locally in one-click via Hermes! ✨ Qwen3.8-27B, Qwen3.8-Flash, DeepSeek-V4-Flash and more are all supported.
Super excited to collab on this!
I’m excited to finally announce the newest edition my Stanford course 𝗧𝗵𝗲 𝗠𝗼𝗱𝗲𝗿𝗻 𝗦𝗼𝗳𝘁𝘄𝗮𝗿𝗲 𝗗𝗲𝘃𝗲𝗹𝗼𝗽𝗲𝗿. It has been 9 months in the making. Last November, with the release of Claude Opus 4.5, coding agents experienced a step function improvement in capability. We all felt it. The LLMs were more powerful, could reason for longer, solve harder tasks. This year’s iteration of my course reflects the 2026 metamorphosis of software engineering. My core belief is simple: AI-native developers of the LLM era are going to become the most important members of any software organization. I have designed my course to train this next generation of engineers. 𝗪𝗵𝗮𝘁’𝘀 𝗱𝗶𝗳𝗳𝗲𝗿𝗲𝗻𝘁 𝘁𝗵𝗶𝘀 𝘁𝗶𝗺𝗲 𝗮𝗿𝗼𝘂𝗻𝗱 First, 85% of my Fall 2025 class material is being thrown out. The Fall 2026 syllabus reflects the core capabilities AI-native engineers must have: agent skills, advanced context engineering, MCP portals, agent-ready codebase principles, agentic code review, security, parallelizing background agents, software factories, and more. Second, I am going to teach my students how to have software taste. Every student will be required to ship pull requests to production-grade, real-world codebases. The course is collaborating with the top open-source AI repos who will offer support and mentorship to students on how to meaningfully contribute to their projects. This has never been done before in any university course so I am incredibly grateful to our OSS Partners: @browserbase, @HeyGen, @CopilotKit, @semgrep, @OpenHandsDev, @milvusio, @marimo_io, Pi, @crewAIInc, @warpdotdev, @vercel, @cmux, @arizeai, @UnslothAI, and @anyscalecompute. 𝗪𝗵𝗮𝘁’𝘀 𝘀𝘁𝗮𝘆𝗶𝗻𝗴 𝘁𝗵𝗲 𝘀𝗮𝗺𝗲 I’m fortunate to again have AI software engineering leaders and founders as guest speakers to share their learnings from building top coding agent products. Thank you to @leerob from @cursor_ai, @bcherny of @claudeai code, @EnoReyes of @FactoryAI, @silasalberti of @cognition, @0xine of @semgrep, Rajesh Bhatia of @Cloudflare , @amasad of @Replit, and @eladgil. All resources will be available online. All classes will be available to the public. 9/22 on Stanford campus. See you in class. 
Show more
Qwen3.8-Flash can now run 1.7× faster locally with MTP!⚡️ GGUFs can reach 170 tokens/s on a RTX PRO 6000. MTP enables Qwen3.8-Flash-Next ~1.3–1.7× faster inference with no accuracy change. GGUFs: Guide:
Show more
Qwen3.8-Flash can now be run locally! 🔥 The 125B MoE model outperforms Claude-Opus-4.6 (Max). Run on 75GB RAM via Unsloth GGUFs. Qwen3.8-Flash-Next enables CPU RAM / unified mem setups to deliver near VRAM speeds. Guide: GGUF:
Show more
0
64
1.2K
129
Forward to community
⚡ Fine-tune and quantize AI models directly on NVIDIA Jetson. Our new Jetson AI Lab tutorial shows how to use Unsloth and memory-efficient QLoRA to customize models, export quantized GGUF files, and run them locally with llama.cpp. Follow hands-on examples for: 🔹 NVIDIA Nemotron 3.5 Lightning on Jetson AGX Thor 🔹 Qwen3.5-4B on Jetson Orin Nano Start optimizing:
Show more
We quantized GLM-5.3 to dynamic 1-bit (217GB) vs BF16 (1.5TB) and it retains ~76% top-1% accuracy whilst being 83% We did a small basic snake game in Unsloth Desktop using the UD 1-bit and it worked well! zai is on a roll with GLM-5.3-Flash and now GLM-5.3!
Show more
GLM-5.3 can now be run locally! The 2-bit model retains ~81% accuracy after we shrunk it from 1.51TB to 239GB (-83% size). Run on a 256GB Mac or RAM/VRAM setups. GLM-5.3 is the strongest open model to date. Guide: GGUF:
Show more
GLM-5.3 can now be run locally! The 2-bit model retains ~81% accuracy after we shrunk it from 1.51TB to 239GB (-83% size). Run on a 256GB Mac or RAM/VRAM setups. GLM-5.3 is the strongest open model to date. Guide: GGUF:
Show more
0
98
1.9K
207
Forward to community