On the SAME 48GB M5 Pro running ๐ง Qwen3.6-35B-A3B @ 210 tps and @ 143 tok/s at 32K context, the multi-agent handling is so much better than mere tokens.
With 32K of context already cached, Splash produced the first token in โก 123 ms vs ๐ 816 ms for oMLX
And when Inco tested 16 simultaneous 32K requests with Qwen3.8-27Bโฆ
Splash handled all 16.
Itโs built to handle heavy agent workloads with lots of context and repeated prompts even with multiple agents at once
w/fast reuse of cached context
Qwen3.8-27B at up to 144 tok/s on an M5 Max MacBook Pro?! You have to try this Splash Engine!
It's a new open-source inference engine called Splash.
Instead of being a universal runtime like llama.cpp or Ollama, Splash optimizes the whole stack around the exact model ๐
โ๏ธ model-specific kernels
๐ง hardware-aware memory planning
๐ DFlash2 speculative decoding
๐พ prompt-cache reuse
๐ฅ continuous batching
On the same 48GB M5 Pro running Qwen3.8-27B:
๐ Splash: 74 tok/s
โก oMLX: 38 tok/s
๐ Ollama: 24 tok/s
At 32K context:
๐ Splash: 54 tok/s
And with 4 concurrent requests:
๐ฅ 170 aggregate tok/s
vs 43 tok/s for oMLX
Now Inco + LM Studio are showing up to ~144 tok/s on M5 Max.
And LM Studio Bionic 1.1.5 already added Splash as an experimental runtime.
Requirements
๐ M3 or newer
๐พ 36GB minimum
๐ 48GB+ recommended
Completely different performance because the software stack is optimized around what it is running.
โ ๏ธ 144 tok/s is an Inco/LM Studio result. The detailed numbers above are from Inco's 48GB M5 Pro test.
๐