I'm done with Bonsai 2 27B. This model has taken enough of my time and patience.
After the whole argument, I ran 12 OMP sessions on my 4x RTX 3090 rig using PrismML's launcher settings and the reasoning budget nisten asked for. 1 session per GPU, 128K context, no follow-up prompts.
Same ternary model, 2 weight packings, 3 KV caches, 2 prompts. Here are the results, the exact prompts, the memory guide and the settings. Do whatever you want with them. Big thread 🧵
First, the extended prompt. 3 of the 6 runs produced usable scenes.
PQ2_0 with Q8 KV made the garden I liked most: 42.0 minutes, 69,240 output tokens.
PQ2_0 with Q4 KV + bias spent 4h54m and 418,141 output tokens producing an empty sky. I stopped it myself. FP16 KV also failed to produce a usable scene after 151.7 minutes and 210,512 output tokens.
The videos show every configuration, including the failures. Output tokens include reasoning; time is the full agent session.
Exact extended prompt:
Design and create a very creative, elaborate, and detailed voxel art scene of a pagoda in a beautiful garden with trees, including some cherry blossoms. Make the scene impressive and varied and use colorful voxels. Use whatever libraries to get this done but make sure I can paste it all into a single HTML file and open it in Chrome.
I'd call the model tolerable for its size. PrismML's marketing still pisses me off. Simple results and the memory guide below.