We compressed Hy4-preview from 1.5TB to ๏ฝ200GiB GGUF and it still works well !
Meet MIX-STQ1_0.The trick isnโt just going low, itโs deciding where: calibration data picks each layerโs bit-width, some down to 1.31-bit STQ1_0, some up to 2.06-bit IQ2_XXS. Same budget, lower error.
Accuracy barely moves vs BF16
๐ MCP Atlas 83.7โ83.2
๐ SWE-Bench multi 82.9โ81.3
๐ MRCR 81.3โ81.1
๐ IFBench 73.5โ72.5
See the details on HF : AngelSlim/Hy4-preview-GGUF
Weights & low-bit GGUFs ๐
#
LLM# #
Quantization# #
llamacpp# #
Hy#