đĨ Thank you, Liquid AI, for this new quant method. Unsloth, Google, and Nvidia aren't the only ones pushing smaller quantizations.
Liquid AI released new Q4_0 GGUFs that were trained to survive 4-bit quantization.
And they're retaining roughly 97% of their BF16 performance.
Liquid calls the technique đ Quantization-Aware Distillation (QAD) đ§
Normal quantization:
1ī¸âŖ Train a model in high precision
2ī¸âŖ Quantize it afterward
3ī¸âŖ Accept some intelligence loss to make it smaller
QAD changed step #
2#
The quantized model is trained while experiencing the errors introduced by 4-bit quantization, with the high-precision model acting as its teacher.
So it learns to compensate for being Q4.
Liquid applied it to their LFM2.5 models:
đ§ LFM2.5-230M
đ§ LFM2.5-350M
đ§ LFM2.5-1.2B
đ§ LFM2.5-2.6B
The resulting Q4 models retain:
⥠230M â 97.1% of BF16 performance
⥠350M â 96.5%
⥠1.2B â 97.4%
⥠2.6B â 96.6%
And these models are TINY:
đĻ 350M â 219MB
đĻ 1.2B â 696MB
đĻ 2.6B â 1.59GB
The 230M and 350M QAD models reportedly reach roughly Q5_K_M quality while keeping the faster Q4_0 inference path.
And they're available right now as GGUFs and the Local AI ecosystem can already run them
đĻ llama.cpp
đĸ Ollama
đĨī¸ LM Studio
đ Lemonade
đ¤ Hermes Agent
đĻ OpenClaw
Run them on:
đą Phones
đģ Laptops
đĨ§ Raspberry Pi
đĨī¸ Mini PCs
And now I really want Liquid to do this to LFM2.5-8B-A1B. đ