bonsai 2 27b just built this from one paragraph of prompt in one shot, all of it out of a 5.9gb file on an rtx 3060 12gb.
i did not expect frontend taste at this size. small models usually get the logic right and the layout wrong, this one got the layout right, and it thought for a long time to do it, 41k tokens over 46 minutes, the context ran out to 77k and it never lost the thread.
this is a ternary compression of qwen 3.8 27b, 26 tok/s fresh on a five year old gaming gpu, 13 tok/s at 77k deep, the whole 262k window resident. for what it is, on this card, this is insane, and i cannot wait to run it through real agentic coding on hermes agent, the tool loop, the builds that break and have to recover.
@PrismML keep going, this is the one that runs on the card people actually own.