Register and share your invite link to earn from video plays and referrals.

Search results for bonsAI倶楽部
bonsAI倶楽部 community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including bonsAI倶楽部
Bonsai 2 has been evaluated with a low thinking budget for xhigh. Quantization errors really show their impact on long sequences, and Qwen3.8 27B often needs more than 81K tokens to complete its answer. For coding problems, like in LiveCodeBench, this is not enough. Expect some surprises for long-horizon agentic tasks. It's probably not as good as the model card says. Remarkable work nonetheless, as always.
Show more
I'm done with Bonsai 2 27B. This model has taken enough of my time and patience. After the whole argument, I ran 12 OMP sessions on my 4x RTX 3090 rig using PrismML's launcher settings and the reasoning budget nisten asked for. 1 session per GPU, 128K context, no follow-up prompts. Same ternary model, 2 weight packings, 3 KV caches, 2 prompts. Here are the results, the exact prompts, the memory guide and the settings. Do whatever you want with them. Big thread 🧵 First, the extended prompt. 3 of the 6 runs produced usable scenes. PQ2_0 with Q8 KV made the garden I liked most: 42.0 minutes, 69,240 output tokens. PQ2_0 with Q4 KV + bias spent 4h54m and 418,141 output tokens producing an empty sky. I stopped it myself. FP16 KV also failed to produce a usable scene after 151.7 minutes and 210,512 output tokens. The videos show every configuration, including the failures. Output tokens include reasoning; time is the full agent session. Exact extended prompt: Design and create a very creative, elaborate, and detailed voxel art scene of a pagoda in a beautiful garden with trees, including some cherry blossoms. Make the scene impressive and varied and use colorful voxels. Use whatever libraries to get this done but make sure I can paste it all into a single HTML file and open it in Chrome. I'd call the model tolerable for its size. PrismML's marketing still pisses me off. Simple results and the memory guide below.
Show more
Try 1-bit Bonsai 27B, the latest model from @PrismML on Mac. This model pushes the limits of intelligence density, packing 27B parameters into just 5 GB. Powered by MLX. Available now.
Show more
Benchmark results for Bonsai Q2_0 and Q1_0 on my GX10, using strict EvalPlus: Ternary Bonsai Q2_0 • HumanEval+: 152/164 92.68% • MBPP+: 305/378 80.69% • Total Plus solved: 457/542 Bonsai Q1_0 • HumanEval+: 149/164 90.85% • MBPP+: 282/378 74.60% • Total Plus solved: 431/542 Both completed all 542 generations with zero errors. Ternary recovered 26 tasks over Q1_0. Of those, 23 came from MBPP+ and only 3 from HumanEval+. For context, PrismML reports Qwen3.6-27B FP16 at 95.12% HumanEval+, 83.33% MBPP+, and 88.74% across its coding category including LiveCodeBench. That's a very very solid model for such weights and considering its file size.
Show more
Meet 1-bit Bonsai 27B, the latest model from @PrismML and the first 27B-class model to run on a phone. This model, available on the iPhone 17 Pro, iPhone Air, and select iPads, pushes the limits of intelligence density. Powered by MLX. Available now.
Show more
Got a new bonsai tree So calming to care for. Plant lovers, what’s your fave? Bonsai 🌻
Try the new Ternary Bonsai 8B from @PrismML. A larger, smarter Bonsai model available on iPhone and iPad. Update your app now.
How did I miss this?! Bonsai isn’t just an LLM family. Back in May, PrismML released Bonsai Image 4B, a crazy low-bit version of FLUX.2 Klein 4B designed to run locally on 🍎 iPhones. Look at this 🎨 FLUX.2 Klein 4B transformer 💾 FP16: 7.75GB 🌳 Ternary Bonsai: 1.21GB And the entire Apple deployment payload, including its compressed text encoder + VAE, is only: 🔥 3.88GB 📱 iPhone 17 Pro Max 🧠 A19 Pro / 12GB unified memory 🖼️ 512×512 ⚡ 9.4 sec/image 📲 MLX Swift ☁️ No cloud 💾 Bomsai Studio (App Store) M4 Pro: ~5.8 sec/image. No cloud. The ~4B diffusion transformer is paired with a 4-bit Qwen3-4B text encoder, which gets unloaded after the prompt is encoded to save memory. And there’s also 🍎 Mac / iPhone / iPad support 🟢 low-bit Gemlite builds for Nvidia GPUs 🔓 Apache 2.0 This is completely separate from the Bonsai 2 27B LLM I’ve been posting about. PrismML took a 7.75GB FLUX transformer and crunched it down to 1.21G and put image generation on an iPhone. 👀 And somehow I missed this for 4 months.
Show more
So I asked ChatGPT to write a headless server for my Debian + 4060Ti 16Gb that loads bonsai 2 27B and provides an interactive shell. It decided to go for the 6Gb version. 5min later... I can't be blasé about this.
Show more
oMLX 0.7.0rc1 is out! This release brings faster Qwen prefill & generation, MiMo V2.6, Ternary Bonsai 2, and partial block caching. DFlash now handles concurrent requests together, and Lightning MTP gets faster batch decoding! Performance on M5 Max, 128 GB (Prefill, oQ4e quant) - Qwen3.8-Flash-Next: 1,522 -> 2,007 tok/s (+32%) at 16K context. (Decode, batch=4, oQ4e quant) - Qwen3.8-27B with DFlash2: 56.9 -> 131.5 tok/s (+131%) - Qwen3.8-27B with Lightning MTP: 88.9 -> 136.9 tok/s (+54%) Full benchmark details are in the release notes. New models and features - MCDMA RDMA support for Mac + CUDA deployments, contributed by @ashxhart. - Partial block caching. No more reprocessing thousands of tokens just because they didn't fill a complete cache block. In one test, next-turn prefill dropped from 1,174 tokens to 37. - Ternary Bonsai 2 text and vision support. - MiMo V2.6 image, video, and audio understanding, plus Lightning MTP and DFlash for compatible checkpoints. - Broader MoE expert offload, including Lightning MTP alongside expert offload for DeepSeek V4.1 and GLM-5.3-Flash. This RC also includes the improvements from the dev releases, including Cluster v2, one-click model settings from community benchmarks, and a customizable dashboard. The GDN prefill kernels are adapted from @ddalcu's excellent mlx-serve! After a short round of testing, I'll publish the stable release and keep moving forward!
Show more