Bonsai 2 27B by
@PrismML is genuinely one of the few local models I can run reliably on my M3 Pro.
The coding and agentic performance is seriously good. It handles long tasks, tool calls, file operations, code execution, and vision/multimodal work without feeling like a stripped-down local model.
The main downside is speed.
I’m getting around 8.5–11 tok/s, but this is a 27B model running at 1-bit locally. Getting this level of capability on-device at that size is pretty impressive.
Today, we’re announcing Ternary Bonsai 2 27B.
Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance.
Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use.
Ternary Bonsai 2 27B is available today under Apache 2.0.
Show more