佔用還是 5.9 GB。兩個月前也是 5.9 GB。
變的是底盤。換成 Qwen3.8 27B,就是我現在桌上那顆。而且是 dense,不是 MoE。每個字出來,權重幾乎整包都在場,不是只點亮一小塊。
上一代壓完大概留全精度的九成五。這代寫 98.2%。箱子一樣大,差的那一截被補回來了。數學、寫 code 幾乎貼著原模型。比較會掉的是看圖。
所以這不是又一個「更小的模型」。是同一顆日常在用的 27B,第一次真的能塞進筆電、消費級顯卡還看起來不像被壓壞。
速度我還沒在自己機器上對過。官方寫 5090 可以到 143 tok/s,M5 Max 大約 47。先當他們的表,不當我的結論。
7 月那則講的是:27B 進得了 4–6 GB,local 開始像產品。這次沒把箱子再壓小。是把「跟原模型差多少」從明顯縮到幾乎貼著。
雲不會消失。特別難的還是會丟上去。但日常那層如果可以留在自己桌上,付的就不是每一個 token 的 API 帳單。
HBM、電、機櫃還是在簽。被打穿的是模型那一層的收費,不是機房那一層的訂單。
Today, we’re announcing Ternary Bonsai 2 27B.
Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance.
Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use.
Ternary Bonsai 2 27B is available today under Apache 2.0.
顯示更多