Register and share your invite link to earn from video plays and referrals.

Sean
@sean_________
TMT catalyst/event-driven trader with a focus on research.
Joined May 2019
3K Following    9.8K Followers
UBS Hot Chips 2026- Model size + KV cache keep going up. Roadmaps are now capacity, bandwidth, and scale-up — not peak FLOPS. $GOOG TPU 8i (inference, 2-die) is a memory-bandwidth chip: 384MB on-die SRAM, 288GB HBM3E, 19.2Tb/s inter-chip. Built because reasoning / MoE / agentic decode is now more memory-bound than training. TPU 8t ( $AVGO, training): 9,600 chips/Superpod, 121 EF FP4, 2PB shared HBM. Different product. Training and inference silicon have diverged enough to justify two architectures. Same conference, same bottleneck everywhere else: Samsung LPDDR5X-PIM: MAC inside the DRAM banks. 614GB/s PIM BW (8x LPDDR5X), ~3x Llama-3.1-8B tokens. Cheap inference vs HBM. CXL pooling (Samsung demo): 3.35x LLM decode at 100K context. d-Matrix 3D-DRAM: META engineer on stage. 32GB/card, ~100TB/s, 1M-token context in a 72-card rack. LPX (Groq/$NVDA): no HBM at all. SRAM-only decode. LPX + Rubin ~3-5x on long-context agentic. $CBRS CS-6: 3D DRAM on wafer-scale SRAM. ~2028. Net: TPU 8i is Google saying inference is a memory-movement problem. PIM / CXL / 3D-DRAM / SRAM decode are everyone else saying the same thing. AVGO still has the training Superpod. The duration is the memory stack sitting under both.
Show more