Register and share your invite link to earn from video plays and referrals.

Search results for GGUF
GGUF community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including GGUF
We released experimental MTP Qwen3.6 Unsloth GGUFs! Qwen3.6 27B MTP now runs at 140 tokens/s. Qwen3.6 35B-A3B MTP gets 220 tokens/s generation on a single GPU. Qwen3.6 27B and 35B-A3B have >1.4x speed-up over the original GGUFs without any change in accuracy. Guide + GGUFs + Benchmarks: In terms of average speedup, we see a 1.4x for dense models at draft tokens = 2 and for the MoE around 1.15 to 1.2x. We do not recommend more than 2 draft tokens because the acceptance rate drops precipitously from 83% to 50% with 4 draft tokens, and the forward passes for MTP become less beneficial. Use `--spec-type mtp --spec-draft-n-max 2` Thanks to Aman for
Show more
0
60
764
113
Forward to community
Minimax H3 😃Video Gen models + workflows (NVFP4/BF16/FP8/INT8/INT4/GGUFs) community released All listed at one place 👇😲
Gemma 4 now runs 2x faster with MTP GGUFs! Run locally on just 6GB RAM. ⚡️ MTP enables Google Gemma 4 run ~1.4–2.2× faster with no accuracy loss. Gemma 4 12B MTP can run at 162 t/s vs. 52 t/s without MTP. 31B reaches 101 t/s. GGUFs + Guide:
Show more
0
62
2.2K
259
Forward to community
Update on the DwarfStar laguna-s2.1 branch, the GGUF files got updated, now the support for the file with Q8 projections and smaller routed experts (68GB total IIRC) is implemented as well. Speed improved significantly to 60 t/s generation, 550 t/s prefill.
Show more
OpenClaw landed on @huggingface local apps 🦞🤝🤗 1. Pick any GGUF/MLX model on hf 2. Copy the openclaw onboard setup 3. Volla you've got a tool-calling agent running fully local. no cloud, no keys, no one watching. Get your claw localmaxxing. resistance is futile 🦞
Show more
Local AI is having its moment! Below is the number of new GGUF models created each month over the past 8 months & insights from our HF internal agent (May is partial): - 176,000 total public GGUF models on HF - Two distinct regimes: Oct–Feb averaged ~5.1K new GGUF models/month. Then March–April jumped to ~9.2K/month — nearly double the previous rate. - March was the inflection point (+55% MoM) — likely driven by a wave of new open-weight model releases being quantized to GGUF. - April sustained the momentum at 9.7K, suggesting this isn't a one-off spike but a new baseline. - The GGUF ecosystem is accelerating — the community is quantizing models faster than ever, likely thanks to better tooling (llama.cpp improvements, automated quantization pipelines, and more models supporting GGUF natively). Let's go!
Show more
DeepSeek-V4-Flash-0731 (all 284B parameters) running on a single DGX Spark. ⚡ ~17 tok/s generation, 40 tok/s prefill 📦 One 80GB GGUF file, mixed IQ2_XXS/Q8 imatrix quant 🧠 256-expert MoE, 6 routed per token 🆓 MIT licensed frontier model (kinda) on your desk 😎 🤗
Show more
Hy-MT2 keeps gaining momentum. Since its open-source release in May: → 700K+ downloads 🌟 → Hy-MT2-1.8B reached #1# on the Hugging Face trending, with 30B-A3B reaching #4# 🥇 → 70+ verified product and project integrations 💻 → Broader Hy-MT ecosystem support across Apple MLX-LM, Microsoft ONNX Runtime, NVIDIA NeMo, LLaMA-Factory, and more 👯 → Real-world adoption, including real-time multilingual translation of livestream comments on Bilibili 📺 And now, Hy-MT2-30B-A3B is officially available in GGUF format—addressing one of the community’s most-requested deployment needs and making local inference easier. Ready to run Hy-MT2-30B-A3B locally? Try the new GGUF release: Explore Hy-MT2: HuggingFace: Modelscope: Github: #TencentHy# #HyMT2# #OpenSource#
Show more