๐๏ธ NEW EPISODE: Tensordyne co-founder and CPO R K Anand joins Austin to break down an AI inference rack built on a router-based scale-up fabric and logarithmic math.
- HPE Juniper's 7th-gen router fabric provides the scale-up network at 1-2 microseconds of latency
- Log math turns expensive multiplications into cheap adds
- The compute die stays below the reticle limit, leaving room for a big on-chip SRAM cache
- 72 chips in a 13U chassis at 30 kW, a quarter of the space and power of an NVL72
- Air cooling opens brownfield enterprise and telco data centers
- One system runs both pre-fill and decode
@TensordyneInc @_rk_anand_
Chapters:
0:00 Introducing Tensordyne
5:32 The Juniper vs. Cisco Playbook
11:29 Origin Story: Automotive Power Constraints
15:37 The Secret Sauce of Log Math
18:02 Pivoting to the Data Center
22:08 Leveraging a Router Backplane for AI
27:22 Why Router Fabrics Suit MoE Models
34:12 The Three Phases of Inference Hardware
37:40 How One Chip Handles Pre-fill & Decode
40:34 The 'Too Good to Be True' System Specs
43:31 Go-to-Market: The Air-Cooled Advantage
48:21 De-risking with Strategic Partnerships
52:37 Solving the Software Problem with AI
Get more of Austin and Vik daily, free!
Sign up:
Connect with Vik and Austin:
Vik's Paid Substack:
Austin's Paid Substack:
@austinsemis @vikramskr