K2 Horizon got another llama.cpp-family path, this time through TurboQuant.
Spark-X2.5 already landed in upstream llama.cpp. I’ve actually got it running myself.
A brand-new K2 Horizon TurboQuant PR adds …
🧠 K2 Horizon model support
📦 Hugging Face → GGUF conversion
💬 tokenizer + chat templates
⚙️ llama.cpp-style inference
And they’ve already tested
🧠 K2-Horizon-3.7B Q4_K_M
🍎 M3 Pro / 18GB
⚡ ~35 tok/s
📚 64K configured context
The PR also adds Spark-X2.5 to the TurboQuant fork:
⚡ ~36 tps
📚 256K configured context
Caveats
⚠️ This is still an OPEN PR
⚠️ Only Metal has been tested
⚠️ Other K2 sizes + MoVA/MoE variants remain untested
K2 Horizon already has an IFM llama.cpp fork, but it still isn’t supported in normal upstream llama.cpp.
So what I’m watching here is whether TurboQuant becomes another path for K2.