đ¤¯ Xiaomi showed a 150W mini AI box designed to run a 120B model LOCALLY.
Yeah, it's Xiaomi, so good luck getting it in the US if it ever ships.
The Xiaomi AI Cube, a prototype built around 3 of Xiaomi's own XRING chips.
And the specs are kind of crazy đ
đ§ Local deployment â 120B + 3B models
đž D100 â up to 160GB unified memory đ
đ O100 â 1.22 TB/s near-memory bandwidth đ Nvidia did you see this?
⥠Entire AI Cube â up to 150W
đ§Ž O3 â 200 TOPS NPU
đŽ O3 â 16-core G2 Ultra NX GPU
That 1.22 TB/s number is especially interesting.
According to Xiaomi, they stack high-speed DRAM directly over the O100's logic/NPU layer using wafer-on-wafer packaging + hybrid bonding.
In other words đ very short path between memory and compute.
And Xiaomi says the D100 itself can accommodate local models as large as 200B parameters. đ
đ¯ Strix Halo showed what 128GB unified memory could do. Gorgon Halo is up next with 192GB.
Now Xiaomi is experimenting with 160GB-class unified memory + >1 TB/s bandwidth in a 150W mini box.
Caveats of course ...
â ī¸ It's a prototype
â ī¸ No price
â ī¸ No retail release date
â ī¸ O100/D100 commercial use is planned for 2027
đ Source: GizmoChina