🔥 Thank you, Liquid AI, for this new quant method. Unsloth, Google, and Nvidia aren't the only ones pushing smaller quantizations.
Liquid AI released new Q4_0 GGUFs that were trained to survive 4-bit quantization.
And they're retaining roughly 97% of their BF16 performance.
Liquid calls the technique 👉 Quantization-Aware Distillation (QAD) 🧠
Normal quantization:
1️⃣ Train a model in high precision
2️⃣ Quantize it afterward
3️⃣ Accept some intelligence loss to make it smaller
QAD changed step #
2#
The quantized model is trained while experiencing the errors introduced by 4-bit quantization, with the high-precision model acting as its teacher.
So it learns to compensate for being Q4.
Liquid applied it to their LFM2.5 models:
🧠 LFM2.5-230M
🧠 LFM2.5-350M
🧠 LFM2.5-1.2B
🧠 LFM2.5-2.6B
The resulting Q4 models retain:
⚡ 230M → 97.1% of BF16 performance
⚡ 350M → 96.5%
⚡ 1.2B → 97.4%
⚡ 2.6B → 96.6%
And these models are TINY:
📦 350M → 219MB
📦 1.2B → 696MB
📦 2.6B → 1.59GB
The 230M and 350M QAD models reportedly reach roughly Q5_K_M quality while keeping the faster Q4_0 inference path.
And they're available right now as GGUFs and the Local AI ecosystem can already run them
🦙 llama.cpp
🟢 Ollama
🖥️ LM Studio
🍋 Lemonade
🤖 Hermes Agent
🦞 OpenClaw
Run them on:
📱 Phones
💻 Laptops
🥧 Raspberry Pi
🖥️ Mini PCs
And now I really want Liquid to do this to LFM2.5-8B-A1B. 👀