Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro ⚡
Meet Inco Splash: our open-source inference engine, built around the model and around Apple silicon.
Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.
⚡ 532 tokens/s!
DeepSeek V4.1 Flash is live on Inco, and it's the fastest provider on Artificial Analysis. #1# output speed. Not close.
Try it:
Results: