K2 Horizon got another llama.cpp-family path, this time through TurboQuant.
Spark-X2.5 already landed in upstream llama.cpp. Iโve actually got it running myself.
A brand-new K2 Horizon TurboQuant PR adds โฆ
๐ง K2 Horizon model support
๐ฆ Hugging Face โ GGUF conversion
๐ฌ tokenizer + chat templates
โ๏ธ llama.cpp-style inference
And theyโve already tested
๐ง K2-Horizon-3.7B Q4_K_M
๐ M3 Pro / 18GB
โก ~35 tok/s
๐ 64K configured context
The PR also adds Spark-X2.5 to the TurboQuant fork:
โก ~36 tps
๐ 256K configured context
Caveats
โ ๏ธ This is still an OPEN PR
โ ๏ธ Only Metal has been tested
โ ๏ธ Other K2 sizes + MoVA/MoE variants remain untested
K2 Horizon already has an IFM llama.cpp fork, but it still isnโt supported in normal upstream llama.cpp.
So what Iโm watching here is whether TurboQuant becomes another path for K2.