How did I miss this?! Bonsai isn’t just an LLM family.
Back in May, PrismML released Bonsai Image 4B, a crazy low-bit version of FLUX.2 Klein 4B designed to run locally on 🍎 iPhones.
Look at this
🎨 FLUX.2 Klein 4B transformer
💾 FP16: 7.75GB
🌳 Ternary Bonsai: 1.21GB
And the entire Apple deployment payload, including its compressed text encoder + VAE, is only:
🔥 3.88GB
📱 iPhone 17 Pro Max
🧠 A19 Pro / 12GB unified memory
🖼️ 512×512
⚡ 9.4 sec/image
📲 MLX Swift
☁️ No cloud
💾 Bomsai Studio (App Store)
M4 Pro: ~5.8 sec/image.
No cloud.
The ~4B diffusion transformer is paired with a 4-bit Qwen3-4B text encoder, which gets unloaded after the prompt is encoded to save memory.
And there’s also
🍎 Mac / iPhone / iPad support
🟢 low-bit Gemlite builds for Nvidia GPUs
🔓 Apache 2.0
This is completely separate from the Bonsai 2 27B LLM I’ve been posting about.
PrismML took a 7.75GB FLUX transformer and crunched it down to 1.21G and put image generation on an iPhone. 👀
And somehow I missed this for 4 months.