Register and share your invite link to earn from video plays and referrals.

Search results for InferenceX
InferenceX community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including InferenceX
AgentX - InferenceXv3: Does CUDA Moat Hold up in Agentic Inferencing? $3 Million USD dataset open sourced, 1 Mil+ Context Length, Multiturn, Sub Agents 95%+ KVCache HitRate, GB300 NVL72, MI355, B200
Show more
ALERT🚨🚨: On apples-to-apples InferenceX Comparisons, TPUv7 Ironwood achieves 50% better perf per dollar than Blackwell Ultra on the new TorchTPU external inference stack! TorchTPU brings native PyTorch to TPUs, & along with Google open-sourcing a bunch of their Pallas inference kernels, Google has laid out a solid foundation for TPU to rapidly externalize. We at SemiAnalysis strongly believe that TPU externalization is heading in the right direction and moving full steam ahead.
Show more
TPU Inference Externalization Full Steam Ahead - InferenceX, Up to 50% Better Performance per Dollar, Rapid Externalization of TPU stack, Growing Customer Base, Ironwood, TPUv8i, Reducing CUDA Moat
Show more
RT @SemiAnalysis_: TPU Inference Externalization Full Steam Ahead InferenceX, Up to 50% Better Performance per Dollar Rapid Externalization…
Proud to see Mooncake featured among the supporters of @SemiAnalysis_ InferenceX initiative. Reliable, transparent benchmarking is critical as inference systems become increasingly disaggregated and complex. We’re excited to contribute to this ecosystem and help push the frontier of efficient, scalable AI inference. More details: #InferenceX# #AIInference# #Mooncake#
Show more
Engrams Embedding Entendre: Codesign for Efficient DRAM/SSD Offloading New Model Architecture Implications for TAM of DRAM/NVMe, DeepSeek V4.1 Flash, AgentX, InferenceX, NVMe experiments
Show more
Vera Rubin NVL72 Agentic Inference: 67x better Performance per Dollar Jensen Sandbagging Performance Again, 2x more Annual Profit Per GigaWatt, The More you Buy, The More you Earn, AgentX, InferenceX, Extreme Co-Design
Show more
Nvidia Chief Scientist Bill Dally: "What everyone is showing, acutally did see some slides on this at Hot Chips, is the SemiAnalysis InferenceX Benchmark. Cause it hits exactly the data that people want" Nvidia Chief Scientist stated this during a Compute History Museum panel with Google TPU Fellow Norm Jouppi last month.
Show more
Nvidia Chief Scientist Bill Dally: "What everyone is showing, acutally did see some slides on this at Hot Chips, is the SemiAnalysis InferenceX Benchmark. Cause it hits exactly the data that people want" Nvidia Chief Scientist stated this during a Compute History Museum panel with Google TPU Fellow Norm Jouppi last month.
Show more
Agentic inference is making KV cache management and data movement increasingly central to the serving stack. The new AgentX / InferenceXv3 analysis from @SemiAnalysis_ captures this shift well. Glad to see Mooncake highlighted for external KV caching and disaggregated P/D data transfer, including support for hybrid-attention cache management and preserving speculative decoding state. Special thanks to AMD and SemiAnalysis for providing the CI resources and technical support that made Mooncake ROCm wheel packaging possible. We’re excited to ship it in our next release. Read more:
Show more