註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | 🔔 Follow for AI & Vibe Coding Tips 👇
加入 July 2023
549 正在關注    11.2K 粉絲
On the SAME 48GB M5 Pro running 🧠 Qwen3.6-35B-A3B @ 210 tps and @ 143 tok/s at 32K context, the multi-agent handling is so much better than mere tokens. With 32K of context already cached, Splash produced the first token in ⚡ 123 ms vs 🐌 816 ms for oMLX And when Inco tested 16 simultaneous 32K requests with Qwen3.8-27B… Splash handled all 16. It’s built to handle heavy agent workloads with lots of context and repeated prompts even with multiple agents at once w/fast reuse of cached context
顯示更多
Qwen3.8-27B at up to 144 tok/s on an M5 Max MacBook Pro?! You have to try this Splash Engine! It's a new open-source inference engine called Splash. Instead of being a universal runtime like llama.cpp or Ollama, Splash optimizes the whole stack around the exact model 👇 ⚙️ model-specific kernels 🧠 hardware-aware memory planning 🚀 DFlash2 speculative decoding 💾 prompt-cache reuse 👥 continuous batching On the same 48GB M5 Pro running Qwen3.8-27B: 🚀 Splash: 74 tok/s ⚡ oMLX: 38 tok/s 🐌 Ollama: 24 tok/s At 32K context: 🚀 Splash: 54 tok/s And with 4 concurrent requests: 🔥 170 aggregate tok/s vs 43 tok/s for oMLX Now Inco + LM Studio are showing up to ~144 tok/s on M5 Max. And LM Studio Bionic 1.1.5 already added Splash as an experimental runtime. Requirements 🍎 M3 or newer 💾 36GB minimum 👍 48GB+ recommended Completely different performance because the software stack is optimized around what it is running. ⚠️ 144 tok/s is an Inco/LM Studio result. The detailed numbers above are from Inco's 48GB M5 Pro test. 🔗
顯示更多