๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
๊ฐ€์ž… July 2023
549 ํŒ”๋กœ์ž‰ ์ค‘    11.2K ํŒฌ
๐Ÿ”ฅ Thank you, Liquid AI, for this new quant method. Unsloth, Google, and Nvidia aren't the only ones pushing smaller quantizations. Liquid AI released new Q4_0 GGUFs that were trained to survive 4-bit quantization. And they're retaining roughly 97% of their BF16 performance. Liquid calls the technique ๐Ÿ‘‰ Quantization-Aware Distillation (QAD) ๐Ÿง  Normal quantization: 1๏ธโƒฃ Train a model in high precision 2๏ธโƒฃ Quantize it afterward 3๏ธโƒฃ Accept some intelligence loss to make it smaller QAD changed step #2# The quantized model is trained while experiencing the errors introduced by 4-bit quantization, with the high-precision model acting as its teacher. So it learns to compensate for being Q4. Liquid applied it to their LFM2.5 models: ๐Ÿง  LFM2.5-230M ๐Ÿง  LFM2.5-350M ๐Ÿง  LFM2.5-1.2B ๐Ÿง  LFM2.5-2.6B The resulting Q4 models retain: โšก 230M โ†’ 97.1% of BF16 performance โšก 350M โ†’ 96.5% โšก 1.2B โ†’ 97.4% โšก 2.6B โ†’ 96.6% And these models are TINY: ๐Ÿ“ฆ 350M โ†’ 219MB ๐Ÿ“ฆ 1.2B โ†’ 696MB ๐Ÿ“ฆ 2.6B โ†’ 1.59GB The 230M and 350M QAD models reportedly reach roughly Q5_K_M quality while keeping the faster Q4_0 inference path. And they're available right now as GGUFs and the Local AI ecosystem can already run them ๐Ÿฆ™ llama.cpp ๐ŸŸข Ollama ๐Ÿ–ฅ๏ธ LM Studio ๐Ÿ‹ Lemonade ๐Ÿค– Hermes Agent ๐Ÿฆž OpenClaw Run them on: ๐Ÿ“ฑ Phones ๐Ÿ’ป Laptops ๐Ÿฅง Raspberry Pi ๐Ÿ–ฅ๏ธ Mini PCs And now I really want Liquid to do this to LFM2.5-8B-A1B. ๐Ÿ‘€
๋” ๋ณด๊ธฐ