Benchmark results for Bonsai Q2_0 and Q1_0 on my GX10, using strict EvalPlus:
Ternary Bonsai Q2_0
• HumanEval+: 152/164 92.68%
• MBPP+: 305/378 80.69%
• Total Plus solved: 457/542
Bonsai Q1_0
• HumanEval+: 149/164 90.85%
• MBPP+: 282/378 74.60%
• Total Plus solved: 431/542
Both completed all 542 generations with zero errors.
Ternary recovered 26 tasks over Q1_0. Of those, 23 came from MBPP+ and only 3 from HumanEval+.
For context, PrismML reports Qwen3.6-27B FP16 at 95.12% HumanEval+, 83.33% MBPP+, and 88.74% across its coding category including LiveCodeBench.
That's a very very solid model for such weights and considering its file size.