Register and share your invite link to earn from video plays and referrals.

SemiAnalysis
@SemiAnalysis_
30 Following    160.9K Followers
What is the cost of 1 F-35 fighter in terms of the # of OG Rubin Ultra 1024GB HBM vs. the new Rubin "Ultra" 192GB HBM?
Here's a 52 min deep-dive video into everything AI & power with @SemiAnalysis_. Lot's of work went into that & I think it's worth a watch 😊 👉AI is running out of Power
Show more
aura battles are taking over spain and latin america, with thousands showing up to watch people compete to see who can farm the most aura should semis have its own?
Why did Intel abandon EMIB for their 2027 Diamond Rapids CPU? Uniform Memory Access. All CPU cores in each CBB can access all 16 memory channels with a single hop. UCIe-S through the substrate was required to route signals underneath the near-IMH to the far-IMH with about 30mm of reach to achieve this. In contrast AMD's Venice is a NUMA design. Half the memory channels are on a remote-IOD that require an extra hop for the cores in the CCD to access vs memory channels in the local-IOD. As a bonus, Intel also saves on packaging cost without having to use EMIB, freeing up advanced packaging capacity for their external foundry customers
Show more
On total tokens per $ TCO, a new AMD MI355x submission beats B300 at lower interactivity ranges on AgentX. Shoutout to vLLM, AMD, and LMCache engineers.
OPENAI<>NVIDIA: It is clearly a good idea for OpenAI to start an ASIC program, spend hundreds of millions on R&D, present at Hot Chips, deploy some stuff, scare NVIDIA, and get some more money and backstops from NVIDIA on the order of billions of dollars. If Jalapeno does well at scale, OpenAI will increase its revenue per MW. If Jalapeno doesn't do well, they will still get a lot of financing and backstops from NVIDIA, as Jalapeno is scaring NVIDIA. @sama wins even if their chip loses.
Show more
Shoutout to the cracked team at @vllm_project that implemented recent agentic workload optimizations. (1/5)🧵
NVIDIA Research 🚀 has produced some great research, like LatentMoE (used in Kimi K3) and GatedDeltaNets (used in Qwen). But for e2e frontier training, NVIDIA's bureaucratic culture has produced embarrassing models like Nemotron3 Ultra. Despite NVIDIA Research having amazing talent, Nemotron3 Ultra, with 550B total params (55B active), is getting mogged by all the Chinese models, including even Qwen3.8 27B parameters, which has ~20x fewer parameters.
Show more
This week Doug (@fabknowledge), Sam (@sharshe02) and Jordan (@jordannanos) discuss the OpenAI vs HuggingFace security incident and our recent article on Neocloud security ahead of ClusterMAX 3.0 Full article: 0:00 Cold Open 0:57 Neocloud Security 4:38 The Hugging Face Hack 11:47 Agent Swarm Behavior 14:47 Obliterated Models 20:37 Security as a Service 24:30 What the Data Shows 33:08 Attacker Asymmetry 40:55 Nothing Ever Happens 45:22 CMAX Audit
Show more
Here's an exclusive first look at Samsung's upcoming Exynos 2700 smartphone SoC on their SF2P process! Samsung doubles the prime core count to two with this generation, following others in the industry. We look forward to taking a closer look at this chip when it arrives in 2027!
Show more
Korea’s Trillion-Dollar Sovereign AI Investment: Nvidia Wins, Hynix Loses Korea hosts a Squid Games, National AI Tournament, the best non-Chinese open source model gets eliminated, why Nvidia needs open source, implications for Hynix and Samsung
Show more
Neoclouds have limited cybersecurity. Next time agents successfully go rouge, they'll try taking over a neocloud to run more copies. This is bad. Thus: neoclouds should greatly strengthen their cybersecurity and every company with strong cyber models should help with that.
Show more
Samtech displayed their Co-Packaged Copper (CPC) set-up in the CPX form factor at Taiwan OCP (See the open CPX MSA here: Bandwidth exits the 6.4T pluggable optical engine via a copper connector interface. There are 128 connector pins per module corresponding to 64 differential pairs (DPs) or 32 lanes of 200G. We expect the CPX form factor to present a compelling opportunity for many copper interconnect companies such as TE, Amphenol and Molex even as we see bigger optics TAM in scale-up networking.
Show more
Amazing work by NVIDIA & Google on implementing auto-activation out of the box for Google's NCCL Plugin for ConnectX-7/8 NCCL. This was implemented due to feedback from SemiAnalysis a couple of quarters ago to massively improve the out-of-the-box experience on the GCP platform! Previously, users would need to mess around with the correct library load paths and env vars to get it set up correctly for optimized performance on NVIDIA ConnectX NICs on GCP NVIDIA GPU machines, but now it is fully automated!
Show more
Cerebras is great, & the people at Cerebras are great & lovely! But since NVIDIA showed off their LPX's amazing perf at Hot Chips & OpenAI Jalapeno performance was released, Cerebras folks have been massively cope-tweeting on Twitter. They should focus on continuously improving their software stack & execution instead of cope-tweeting.
Show more
NVIDIA LPU supports 3 types of disaggregated inferencing: 1. Rubin Prefill + LPU Decode for the fastest interactivity 2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN for the middle of the curve 3. Rubin Prefill + Rubin Decode Verification + LPU Drafter for the middle-left of the curve For low interactivity, raw Rubin still takes the win. Looking forward to seeing Rubin + LPU performance curves on open-source agentic benchmarks like AgentX.
Show more
PSMC is the qualified supplier of EMIB's SiCap. Its 12" SiCap capacity plan has been revised upward from 3K wpm by end-2026 to 8-10K wpm in 2H27, with capacity already in place. The line is built on a mature DRAM process and targets EMIB packaging. Meanwhile, 8" SiCap ships 1-2K wpm today, heading to 5K. Our Foundry Industry Model tracks quarterly revenue, capacity, utilization, ASP, demand drivers, and capex across every major foundry, with 2Q26 actuals and a 3Q26 outlook. Reach out to: sales@semianalysis.com
Show more
Ep. 027 - OpenAI Jalapeño: Better Than Nvidia Blackwell (Accelerators) This week Bryan, Myron and Jordan (@JordanNanos) discuss our recent article on OpenAI Jalapeño. They cover the performance, architecture, programming model, implications for NVIDIA and more. 0:00 Cold Open 1:05 Jalapeno Overview 4:22 Tokens Per Megawatt 9:15 Benchmark Caveats 13:48 The CUDA Moat 21:12 How OpenAI Did It 32:13 Samsung HBM4 42:10 AI-Designed Silicon 49:24 Architecture Deep Dive 56:58 Doom and Wrap
Show more