Bonsai 2 has been evaluated with a low thinking budget for xhigh.
Quantization errors really show their impact on long sequences, and Qwen3.8 27B often needs more than 81K tokens to complete its answer.
For coding problems, like in LiveCodeBench, this is not enough.
Expect some surprises for long-horizon agentic tasks. It's probably not as good as the model card says.
Remarkable work nonetheless, as always.