๐ฅ Thank you, Liquid AI, for this new quant method. Unsloth, Google, and Nvidia aren't the only ones pushing smaller quantizations.
Liquid AI released new Q4_0 GGUFs that were trained to survive 4-bit quantization.
And they're retaining roughly 97% of their BF16 performance.
Liquid calls the technique ๐ Quantization-Aware Distillation (QAD) ๐ง
Normal quantization:
1๏ธโฃ Train a model in high precision
2๏ธโฃ Quantize it afterward
3๏ธโฃ Accept some intelligence loss to make it smaller
QAD changed step #
2#
The quantized model is trained while experiencing the errors introduced by 4-bit quantization, with the high-precision model acting as its teacher.
So it learns to compensate for being Q4.
Liquid applied it to their LFM2.5 models:
๐ง LFM2.5-230M
๐ง LFM2.5-350M
๐ง LFM2.5-1.2B
๐ง LFM2.5-2.6B
The resulting Q4 models retain:
โก 230M โ 97.1% of BF16 performance
โก 350M โ 96.5%
โก 1.2B โ 97.4%
โก 2.6B โ 96.6%
And these models are TINY:
๐ฆ 350M โ 219MB
๐ฆ 1.2B โ 696MB
๐ฆ 2.6B โ 1.59GB
The 230M and 350M QAD models reportedly reach roughly Q5_K_M quality while keeping the faster Q4_0 inference path.
And they're available right now as GGUFs and the Local AI ecosystem can already run them
๐ฆ llama.cpp
๐ข Ollama
๐ฅ๏ธ LM Studio
๐ Lemonade
๐ค Hermes Agent
๐ฆ OpenClaw
Run them on:
๐ฑ Phones
๐ป Laptops
๐ฅง Raspberry Pi
๐ฅ๏ธ Mini PCs
And now I really want Liquid to do this to LFM2.5-8B-A1B. ๐