๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
๊ฐ€์ž… July 2023
549 ํŒ”๋กœ์ž‰ ์ค‘    11.2K ํŒฌ
K2 Horizon got another llama.cpp-family path, this time through TurboQuant. Spark-X2.5 already landed in upstream llama.cpp. Iโ€™ve actually got it running myself. A brand-new K2 Horizon TurboQuant PR adds โ€ฆ ๐Ÿง  K2 Horizon model support ๐Ÿ“ฆ Hugging Face โ†’ GGUF conversion ๐Ÿ’ฌ tokenizer + chat templates โš™๏ธ llama.cpp-style inference And theyโ€™ve already tested ๐Ÿง  K2-Horizon-3.7B Q4_K_M ๐ŸŽ M3 Pro / 18GB โšก ~35 tok/s ๐Ÿ“š 64K configured context The PR also adds Spark-X2.5 to the TurboQuant fork: โšก ~36 tps ๐Ÿ“š 256K configured context Caveats โš ๏ธ This is still an OPEN PR โš ๏ธ Only Metal has been tested โš ๏ธ Other K2 sizes + MoVA/MoE variants remain untested K2 Horizon already has an IFM llama.cpp fork, but it still isnโ€™t supported in normal upstream llama.cpp. So what Iโ€™m watching here is whether TurboQuant becomes another path for K2.
๋” ๋ณด๊ธฐ