Bonsai 27B just changed the local LLM game forever.
1-bit quantization shrinks it from 54GB to just 3.8GB (-93%), while retaining 90% of its intelligence. That's insane.
With custom WebGPU kernels written by Fable 5 and GPT 5.6 Sol, the model now runs locally in your browser!
While we eagerly await Fable 5's return, our agentic WebGPU kernel optimization framework kept running.
Opus 4.8 picked up where Fable left off, pushing Liquid AI's new LFM2.5 230M to an unbelievable 1,400 tok/s... running locally in your browser.
Don't blink or you'll miss it.
I think Reachy is the one who needs chess lessons… 😅
Robotics meets WebAI: Gemma 4 running fully offline on WebGPU with Transformers.js, controlling Reachy Mini over WebSerial.
No internet, just a browser and a USB-C cable.
What should Reachy play next?