Register and share your invite link to earn from video plays and referrals.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | 🔔 Follow for AI & Vibe Coding Tips 👇
Joined July 2023
549 Following    11.2K Followers
đŸ”Ĩ Thank you, Liquid AI, for this new quant method. Unsloth, Google, and Nvidia aren't the only ones pushing smaller quantizations. Liquid AI released new Q4_0 GGUFs that were trained to survive 4-bit quantization. And they're retaining roughly 97% of their BF16 performance. Liquid calls the technique 👉 Quantization-Aware Distillation (QAD) 🧠 Normal quantization: 1ī¸âƒŖ Train a model in high precision 2ī¸âƒŖ Quantize it afterward 3ī¸âƒŖ Accept some intelligence loss to make it smaller QAD changed step #2# The quantized model is trained while experiencing the errors introduced by 4-bit quantization, with the high-precision model acting as its teacher. So it learns to compensate for being Q4. Liquid applied it to their LFM2.5 models: 🧠 LFM2.5-230M 🧠 LFM2.5-350M 🧠 LFM2.5-1.2B 🧠 LFM2.5-2.6B The resulting Q4 models retain: ⚡ 230M → 97.1% of BF16 performance ⚡ 350M → 96.5% ⚡ 1.2B → 97.4% ⚡ 2.6B → 96.6% And these models are TINY: đŸ“Ļ 350M → 219MB đŸ“Ļ 1.2B → 696MB đŸ“Ļ 2.6B → 1.59GB The 230M and 350M QAD models reportedly reach roughly Q5_K_M quality while keeping the faster Q4_0 inference path. And they're available right now as GGUFs and the Local AI ecosystem can already run them đŸĻ™ llama.cpp đŸŸĸ Ollama đŸ–Ĩī¸ LM Studio 🍋 Lemonade 🤖 Hermes Agent đŸĻž OpenClaw Run them on: 📱 Phones đŸ’ģ Laptops đŸĨ§ Raspberry Pi đŸ–Ĩī¸ Mini PCs And now I really want Liquid to do this to LFM2.5-8B-A1B. 👀
Show more