Crazy stuff. This caught my eye in the technical report...
Disaggregated inference between Thermodynamic Sampling Units (TSUs) and XPUs / FPGAs.
As I understand it, the TSU is extremely energy efficient at the sparse neural computations this new sparse transformer they've built does so running all of those on the TSU and everything else on an XPU or FPGA is much more energy efficient. It's also faster.