็™ป้Œฒใ—ใฆๆ‹›ๅพ…ใƒชใƒณใ‚ฏใ‚’ๅ…ฑๆœ‰ใ™ใ‚‹ใจใ€ๅ‹•็”ปๅ†็”Ÿๅ ฑ้…ฌใจ็ดนไป‹ๅ ฑ้…ฌใ‚’็ฒๅพ—ใงใใพใ™ใ€‚

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
ๅ‚ๅŠ  July 2023
549 ใƒ•ใ‚ฉใƒญใƒผไธญ    11.2K ใƒ•ใ‚กใƒณ
Qwen3.8-27B at up to 144 tok/s on an M5 Max MacBook Pro?! You have to try this Splash Engine! It's a new open-source inference engine called Splash. Instead of being a universal runtime like llama.cpp or Ollama, Splash optimizes the whole stack around the exact model ๐Ÿ‘‡ โš™๏ธ model-specific kernels ๐Ÿง  hardware-aware memory planning ๐Ÿš€ DFlash2 speculative decoding ๐Ÿ’พ prompt-cache reuse ๐Ÿ‘ฅ continuous batching On the same 48GB M5 Pro running Qwen3.8-27B: ๐Ÿš€ Splash: 74 tok/s โšก oMLX: 38 tok/s ๐ŸŒ Ollama: 24 tok/s At 32K context: ๐Ÿš€ Splash: 54 tok/s And with 4 concurrent requests: ๐Ÿ”ฅ 170 aggregate tok/s vs 43 tok/s for oMLX Now Inco + LM Studio are showing up to ~144 tok/s on M5 Max. And LM Studio Bionic 1.1.5 already added Splash as an experimental runtime. Requirements ๐ŸŽ M3 or newer ๐Ÿ’พ 36GB minimum ๐Ÿ‘ 48GB+ recommended Completely different performance because the software stack is optimized around what it is running. โš ๏ธ 144 tok/s is an Inco/LM Studio result. The detailed numbers above are from Inco's 48GB M5 Pro test. ๐Ÿ”—
ใ‚‚ใฃใจ่ฆ‹ใ‚‹