๐คฏ Xiaomi showed a 150W mini AI box designed to run a 120B model LOCALLY.
Yeah, it's Xiaomi, so good luck getting it in the US if it ever ships.
The Xiaomi AI Cube, a prototype built around 3 of Xiaomi's own XRING chips.
And the specs are kind of crazy ๐
๐ง Local deployment โ 120B + 3B models
๐พ D100 โ up to 160GB unified memory ๐
๐ O100 โ 1.22 TB/s near-memory bandwidth ๐ Nvidia did you see this?
โก Entire AI Cube โ up to 150W
๐งฎ O3 โ 200 TOPS NPU
๐ฎ O3 โ 16-core G2 Ultra NX GPU
That 1.22 TB/s number is especially interesting.
According to Xiaomi, they stack high-speed DRAM directly over the O100's logic/NPU layer using wafer-on-wafer packaging + hybrid bonding.
In other words ๐ very short path between memory and compute.
And Xiaomi says the D100 itself can accommodate local models as large as 200B parameters. ๐
๐ฏ Strix Halo showed what 128GB unified memory could do. Gorgon Halo is up next with 192GB.
Now Xiaomi is experimenting with 160GB-class unified memory + >1 TB/s bandwidth in a 150W mini box.
Caveats of course ...
โ ๏ธ It's a prototype
โ ๏ธ No price
โ ๏ธ No retail release date
โ ๏ธ O100/D100 commercial use is planned for 2027
๐ Source: GizmoChina