Hot Chips 2026 was not just a week of new chip launches. AI inference is being unbundled.
Workload → compute → memory → rack.
This piece is a handbook for understanding why that shift is pushing inference toward heterogeneous computing.
The future of inference is not one faster chip. It is a better division of labor.