Register and share your invite link to earn from video plays and referrals.

Gavin Baker
@GavinSBaker
Managing Partner & CIO, @atreidesmgmt. Husband, @l3eckyy. No investment advice, views my own.
Joined July 2011
6.3K Following    342.7K Followers
Hmm. I might ask @dylan522p if Attention-FFN disaggregation (AFD) is compatible with data locality and not moving the KV cache. That constraint is what OpenAI focused on with Jalapeño. They absolutely did not make an “explicit bet that disagg is not the way to go.” There are multiple forms of disaggregation, and one disaggregated topology compatible with data locality is prefill and attention computed on Jalapeño and FFN on another chip.   And they are obviously doing relatively crude disaggregated inference (PD) at scale today with their GPU fleet. Today.   AFD is a hard networking problem that everyone is working on as it can be generally superior to simpler PD disaggregation via approaching similar interactivity without making the traditional tradeoffs between latency and throughput for growing KV caches.   The CS-4 Wafer I/O Module was described as a “programmable, universal disaggregation interface” designed to support multiple forms of disaggregation, including both prefill-decode disaggregation and attention-FFN disaggregation.   We know that CS-4 is capable of crude PD disaggregation today with GPUs (Helios) and the Trainiums. Likely already generating tokens today in one - and probably both - of these setups.   And I think this is about as explicit a statement as one will get from a public company about Jalapeño and its potential FFN disaggregation companion chip. “Jalapeño+Cerebras.”   And if can do AFD disaggregation with Jalapeño then can almost certainly do it with Trainiums and at a minimum Helios GPUs. Note they probs have technical line of sight to this - would be surprised if they have it working today for all models. And might also be possible to pair Jalapeño with LP30s in an AFD setup.   The more interesting question is what cruder PD disaggregation with an SRAM rack (whether LPX or CS-4) unlocks for the existing GPU fleet. Even Hoppers. And what this might mean for the useful life of Hoppers…
Show more