登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Onur Solmaz
@onusoz
Agents @huggingface 🤗, Maintainer @openclaw 🦞. Experimenting in Making local models work great in OpenClaw. (Spicy) opinions are my own
参加 December 2022
2.2K フォロー中    10K ファン
Speed on Bonsai 2 27b is interesting, especially in relation to the the MoE Qwen3.6-35b-a3b and its anticipated 3.8 version There is a size/speed tradeoff If you have 64 gb of memory for models, then it will probably be better to use the future Qwen 3.8 a3b because it will be faster, around 2x if it uses the same architecture But if you prioritize quality over speed TODAY, then Bonsai 2 27b is a better alternative because it is reportedly a high fidelity compression that would outperform a3b But if you have 32 to 16gb of memory, then rejoice! Your local model just got a huge bump in intelligence! Good day for local models 4 bit quants of a3b are around 20gb in weights, whereas Bonsai 2 27b around 6-7gb, depending on which bit packing you use Don't be fooled by 2 files in the repo. PTQ1_0 and PQ2_0 are the same model, just different ways to do bit packing If you use the smaller bit packing in memory, the PTQ1_0, you need to do more arithmetic operations. That is why prefill PTQ1_0 one is much smaller. But for decode, it's still the memory bytes that are the bottleneck, so the decode with PQ2_0 is lower Let me know if I'm missing something More details on my local profiling on the DGX Spark:
もっと見る