註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | 🔔 Follow for AI & Vibe Coding Tips 👇
加入 July 2023
549 正在關注    11.2K 粉絲
🔥 Thank you, Liquid AI, for this new quant method. Unsloth, Google, and Nvidia aren't the only ones pushing smaller quantizations. Liquid AI released new Q4_0 GGUFs that were trained to survive 4-bit quantization. And they're retaining roughly 97% of their BF16 performance. Liquid calls the technique 👉 Quantization-Aware Distillation (QAD) 🧠 Normal quantization: 1️⃣ Train a model in high precision 2️⃣ Quantize it afterward 3️⃣ Accept some intelligence loss to make it smaller QAD changed step #2# The quantized model is trained while experiencing the errors introduced by 4-bit quantization, with the high-precision model acting as its teacher. So it learns to compensate for being Q4. Liquid applied it to their LFM2.5 models: 🧠 LFM2.5-230M 🧠 LFM2.5-350M 🧠 LFM2.5-1.2B 🧠 LFM2.5-2.6B The resulting Q4 models retain: ⚡ 230M → 97.1% of BF16 performance ⚡ 350M → 96.5% ⚡ 1.2B → 97.4% ⚡ 2.6B → 96.6% And these models are TINY: 📦 350M → 219MB 📦 1.2B → 696MB 📦 2.6B → 1.59GB The 230M and 350M QAD models reportedly reach roughly Q5_K_M quality while keeping the faster Q4_0 inference path. And they're available right now as GGUFs and the Local AI ecosystem can already run them 🦙 llama.cpp 🟢 Ollama 🖥️ LM Studio 🍋 Lemonade 🤖 Hermes Agent 🦞 OpenClaw Run them on: 📱 Phones 💻 Laptops 🥧 Raspberry Pi 🖥️ Mini PCs And now I really want Liquid to do this to LFM2.5-8B-A1B. 👀
顯示更多
0
3
215
30
轉發到社區