Wait? How the hell did Tencent compress Hy4 preview from 1.5TB down to 200GB and barely move the benchmarks??
They basically let calibration data decide how aggressively each layer can be compressed. Some layers got pushed all the way down to 1.31 bit, while the more sensitive ones stay closer to 2 bit so the model doesn’t fall apart. Literally insane efficiency gains.
And somehow MCP Atlas only moves down to 83.2, and SWE-Bench Multi - 82.9 to 81.3.
The low-bit inference progress happening right now is kind of insane.