Register and share your invite link to earn from video plays and referrals.

Search results for AIHardware
AIHardware community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including AIHardware
🧠 Hot Chips 2026: AI's Memory Wall Is Becoming a Money and Packaging Wall Hot Chips 2026 is still producing the usual charts of larger chips and higher peak compute. But this year, memory vendors are becoming central to the AI hardware story. The reason is simple: adding more arithmetic is getting easier. Keeping those units fed is not. Zhihu contributor 乱序摸鱼 reviewed the conference's memory talks and traced a larger shift across Samsung, Micron, SK hynix, and emerging architectures such as HBF. His main takeaway: the boundary between compute and memory is starting to dissolve. 1️⃣ The Memory Wall has become a Money Wall AI is not merely consuming more memory. It is pushing the DRAM industry toward a far more silicon-intensive product. For the same capacity, HBM requires roughly 3× the DRAM die area of DDR. A wafer redirected to HBM can therefore deliver fewer total bits, even before packaging and yield are considered. At the same time, new DRAM capacity can take more than two years to become productive. AI accelerator demand changes much faster. The article cites a revealing mismatch: AI compute can grow around 3× every two years, while HBM bandwidth grows by less than 2× over the same period. HBM solves the bandwidth problem through massive parallelism: wider interfaces, more channels, more banks, TSVs, and advanced packaging. But the bandwidth is not free. Its cost appears elsewhere as silicon area, power, packaging complexity, and yield risk. 2️⃣ The HBM base die is turning into a SoC Traditional HBM had a clean division of labor. The upper DRAM dies stored data, while the base die handled interfaces, testing, and basic routing. HBM4 changes the economics. Once the base die moves to an advanced logic process, using it as an expensive wiring layer becomes increasingly difficult to justify. Samsung's roadmap therefore expands its role in stages. The shorter-distance PHY comes first. Then the memory controller moves down from the xPU, followed by telemetry, repair, testing, RAS, and potentially lightweight processing. This turns HBM from a passive device into something that understands its own banks, temperature, failure state, and access scheduling. But integration creates new tradeoffs. A shorter interface may reduce energy per bit while concentrating the same throughput into a smaller area, raising local power density and creating thermal hotspots. The base die is becoming smarter, but every new function brings more RTL, validation, thermal design, and software work. 3️⃣ Near-memory compute should be measured in bits avoided Putting compute near memory sounds attractive, but available silicon is not enough to justify doing it. The more useful question is: How much data movement disappears after this operator moves closer to memory? Filtering, compression, copying, search, and simple reductions can make sense near memory. They may scan a large dataset and return only a small result, eliminating enormous amounts of traffic. Dense matrix multiplication is different. Data moved into a GPU may be reused many times by Tensor Cores. Moving that computation into a constrained base die could create a weaker accelerator while adding a new compiler and runtime target. The author's rule is practical: move operations that significantly reduce data volume, not operations that merely fit into spare logic area. Samsung's zHBM takes this idea further by vertically stacking HBM above the xPU. The modeled design removes much of the millimeter-scale horizontal path between memory and compute. Samsung's public targets suggest around 70% lower I/O power and roughly 100 W saved in a modeled 1,200 W GPU system. Those are architectural targets rather than proven production results. The shorter electrical path also introduces harder problems in bonding yield, thermal design, power delivery, repair, and cross-die ownership. 4️⃣ HBF could add a cold tier, but software must make it work High Bandwidth Flash proposes another route: use stacked NAND to provide much more capacity at a lower cost than HBM. The tradeoff is severe. HBF may offer cheap capacity, but its bandwidth per gigabyte is roughly an order of magnitude lower. That makes it unsuitable as a drop-in replacement for HBM. Its strongest use cases are large datasets with consistently low access rates: 🔹 Cold MoE experts 🔹 Sparse KV cache 🔹 Prefix cache 🔹 Model state that must remain nearby but is not read every token Even MoE is not automatically a good fit. A single token activates few experts, but a larger batch combines expert requests from many users. The supposedly cold expert pool can become hot surprisingly quickly. The harder problem is system software. An HBM-HBF hierarchy needs placement, allocation, prefetching, request coalescing, cache policy, wear management, and runtime telemetry. The article's verdict is cautious: the architectural need is real, but HBF's path from an attractive model to a dependable product remains largely unproven. 5️⃣ HBM scaling is becoming a packaging problem SK hynix's presentation shows why adding more DRAM layers is no longer a simple capacity upgrade. Moving from 12Hi to 16Hi increases the number of layers by 33%, while the package height rises from roughly 720 μm to 775 μm, an increase of less than 8%. The remaining option is to compress everything. Dies become thinner, inter-die gaps narrower, and bump pitches denser. Warpage, underfill, thermal resistance, and bonding yield all become harder to control. Yield is also cumulative. A defect introduced early may only appear after many expensive processing and stacking steps have already been completed. For 20Hi and beyond, hybrid bonding becomes less of a futuristic option and more of a practical necessity. By removing conventional bumps and thick underfill gaps, hybrid bonding can reduce interconnect pitch and thermal resistance. It can also return part of the height budget to the DRAM itself. SK hynix estimates that, within the same overall stack height, core dies could be up to 24% thicker than with MR-MUF while using an interconnect pitch below 18 μm. That extra silicon thickness matters for mechanical strength, wafer handling, warpage, and manufacturability. 6️⃣ The real boundary being redesigned The most important Hot Chips 2026 memory story is not another increase in TB/s. It is the rising cost of the physical distance between data and compute. As arithmetic moves from FP16 to FP8 and FP4, chips can contain more MAC units than the system can consistently feed. Performance is increasingly lost to memory access, synchronization, interconnect power, thermal limits, and data placement. The optimization target is therefore expanding from one chip to the entire task path. HBM base dies are becoming logic devices. Memory controllers are moving closer to DRAM. Flash may become another managed AI memory tier. Compute and memory may eventually be vertically integrated. None of this eliminates complexity. It decides where that complexity should live so the whole system becomes more efficient. The author's final insight is a good one: Sometimes the best architecture does not make the road faster. It discovers that the trip never needed to happen. 🔗 Full analysis: #HotChips2026# #HBM# #AIInfrastructure# #Semiconductors# #MemorySystems# #AdvancedPackaging# #AIHardware#
Show more
AI demand is becoming a key driver of China's economy: China's chip exports are up to a record ~$290 billion over the last 12 months. This figure has more than doubled over the last 3 years. At the same time, computer and component shipments are up to ~$240 billion, the highest since 2022. Power equipment exports are up to a record ~$190 billion, with the growth accelerating over the last 2 years. This brings China’s total AI-related exports up to ~$720 billion over the last 12 months, an all-time high. Western demand for China-made AI hardware is surging.
Show more
0
89
1.5K
227
Forward to community
AI Hardware Demand Growth and Representative US-Listed Companies June 2026 Executive Summary Nvidia’s transition to the Vera Rubin (VR200) platform marks a significant escalation in AI infrastructure complexity and cost. Our BOM teardown of the next-generation Rubin rack reveals a ~2x increase in total rack cost to approximately $7.8 million (vs. ~$4 million for GB300), driven not solely by the GPU/CPU but by sharp revaluations across the supply chain. Key highlights from downstream components include: • PCB content value +233% YoY, the largest increase. • MLCC +182%, reflecting higher density and count (e.g., ~600k MLCCs per VR200 NVL72 server, +30%+ vs. GB300). • ABF substrates +82%, power solutions +32%, and liquid cooling +12%. These upgrades align with broader AI scaling: 800G/1.6T optical transceivers ramping aggressively, glass-based technologies advancing for packaging and interconnects, and hyperscalers prioritizing performance, power efficiency, and thermal management. We expect sustained multi-year tailwinds for the AI hardware ecosystem into 2027+, with Rubin-driven demand accelerating in H2 2026. Investment Thesis: While Nvidia (NVDA) remains the core beneficiary, the supply chain offers diversified exposure. We favor companies with direct exposure to high-growth areas like advanced PCBs, high-speed optics, and glass substrates/optical interconnects. Risks include execution on new capacity, potential margin pressure from rapid scaling, and geopolitical supply chain factors. 1. PCB: Sharpest Value Uplift in Rubin BOM Morgan Stanley’s detailed analysis shows PCB content in the Rubin rack surging +233% versus GB300. This reflects needs for higher layer counts, advanced materials, better signal integrity, and larger formats to support increased power and interconnect density in AI servers. US Representative: TTM Technologies (TTMI) – Leading US PCB manufacturer with strong positioning in high-complexity boards for data center/AI applications. TTM has invested in capacity expansions (e.g., new facilities) to capture AI-driven demand for advanced HDI and high-layer PCBs. 2. MLCC: Density-Driven Surge Nvidia’s VR200 NVL72 platform requires ~600,000 MLCCs per server, over 30% more than GB300. Combined with the +182% value increase in the BOM, this underscores tightening supply for high-capacitance, high-reliability MLCCs in power delivery and decoupling for AI accelerators. Exposure Note: The MLCC market is dominated by Asian players (e.g., Murata, Samsung Electro-Mechanics, Yageo). US-listed indirect exposure may come through broader electronics or power solution providers, but direct pure-play opportunities are limited. Watch for capacity utilization tightness benefiting the ecosystem. 3. Optical Communication: 800G/1.6T Ramp Accelerating Chinese leader Zhongji Innolight reported Q1 2026 net profit +262% YoY, driven by strong 800G/1.6T shipments, with expectations of significant full-year growth. This mirrors industry-wide momentum as AI clusters shift toward higher-speed optics for reduced latency and power in scale-out/scale-up networking. Nvidia’s investments in photonics and CPO further validate the trend. US Representatives: • Coherent (COHR) and Lumentum (LITE): Key players in optical components and transceivers; Nvidia has made substantial equity investments to secure capacity. • Corning (GLW): Major beneficiary via optical fiber, connectivity, and glass technologies (detailed below). 4. Micro-LED/Glass Substrates & Optical Interconnects: Strategic Partnerships Accelerating On May 20, 2026, BOE announced a cooperation MOU with Corning covering glass-based encapsulation carriers, foldable glass, perovskite substrates, and optical interconnect applications. This aligns with industry shifts toward glass cores for superior flatness, thermal stability, and integration in advanced packaging and photonics—critical for next-gen AI as organic substrates hit limits. US Representative: Corning (GLW) – Central to Nvidia’s optical strategy with multi-billion partnerships, new US optical factories, and expansion in fiber/photonics for AI data centers. Recent deals position GLW for 10x+ capacity growth in key areas. AI Hardware Demand Growth & US-Listed Representative Companies Table Component Demand Growth (vs. GB300) Key Drivers US-Listed Reps Investment Rationale PCB +233% value Higher layers, HDI, signal integrity TTM Technologies (TTMI) Direct AI server/backplane exposure; US capacity expansion MLCC +182% value; +30%+ count Power density in servers Limited direct (ecosystem via power suppliers) Supply tightness supports pricing/volume Optical Comm (800G/1.6T) Strong ramp (e.g., +262% profit ex.) Scale-out networking, CPO transition Coherent (COHR), Lumentum (LITE), Corning (GLW) Nvidia investments; transceiver/fiber boom Glass Substrates/Interconnects Emerging (MOU-driven) Packaging, photonics, thermal/optical Corning (GLW) Nvidia factory deals; US manufacturing tailwinds Power & Liquid Cooling +32% / +12% Higher TDP (e.g., 2300W GPUs) Indirect (ecosystem) Secondary but critical for rack deployment Source: Morgan Stanley BOM analysis, company reports, industry data. Growth metrics approximate from Rubin teardown. Outlook & Risks We project robust 2026-2027 growth in AI capex, with Rubin shipments catalyzing another leg-up in component demand. Optical and advanced substrate shifts could extend the cycle beyond traditional GPU focus. Hyperscalers’ vertical integration and US onshoring (e.g., Corning/Nvidia factories) add resilience. Key Risks: Cyclical capex pauses, yield/execution challenges on new tech (glass/CPO), commodity volatility in passives, and intense competition in Asia-heavy segments. Valuation multiples in the space have expanded; selectivity is key. Recommendation: Overweight select supply chain names with strong Nvidia alignment (e.g., TTMI for PCBs, COHR/LITE/GLW for optics/glass). Monitor Q2 2026 earnings for confirmation of Rubin ramp momentum.
Show more
The AI hardware stack produces more filings, releases, and technical updates than I can manually triage in real time. Most feeds tell me what was published. I wanted a desk that remembers what came before and flags what changed. So I’m experimenting with Grok Bot: monitoring public filings, IR, and conference programs, comparing new disclosures against the previous record, and flagging changes worth a closer look. Not a trading bot. A research desk organized around the AI hardware stack, not around trades. It surfaces the lead. I verify the evidence and decide whether it matters. This was the first overnight run. It surfaced two kinds of change. 1. Credo Revenue finished $4M above the top of guide. GAAP gross margin missed guidance, while non-GAAP remained within it. The more important signal is Q2. The non-GAAP gross-margin range stayed unchanged, while the GAAP range stepped down. The GAAP/non-GAAP gap is now embedded in guidance, not confined to Q1. The Q1 reconciliation shows that the entire 3.5-point gap comes from acquired-intangible amortization ($11.0M) and share-based compensation ($5.7M). Sources: 2. Vertiv The other catch was a new filing rather than a scheduled earnings diff. Vertiv disclosed an agreement to acquire UtilityInnovation Group. The company says the deal extends it upstream “to the grid interconnect, adding microgrid controls, onsite generation and energy storage orchestration, and behind-the-meter power architecture.” My read: that places Vertiv earlier in the data-center power decision chain. Source: It finds. I verify. I post. Early experiment. I’ll share more as I learn what works.
Show more
Before AI: - hardware is hard - teams > solo founders - technical founders only - PhDs build tech in search of a problem - hot SaaS startups get 100x ARR multiples After AI: - hardware is the only thing AI can't build - solo founders are 100x engineers - anyone can build - PhDs building AI infra name their price - SaaSpocalypse Am I missing anything?
Show more
0
176
1.2K
83
Forward to community
The AI you used this morning is still seeing the world through a keyboard and a screen. Nathan Xu @Nathan_Xu_Gao from @PLAUDAI on why conversation is the rawest form of intelligence and what that means for how we build AI hardware.
Show more
As AI and physical devices converge, Qwen is fueling the rise of physical AI across industries, unlocking smarter, more capable real‑world applications. Watch the video to see how it’s powering the next generation of AI hardware.⚙️
Show more
the AI hardware companion of the future wont be a new device that sits on your face or hangs around your neck it will be a combination of existing devices evolved, primarily airpods + apple watch the beauty in this approach is the progressive increase in how "invasive" the devices are to the experience 1. normal life: airpods (with cameras) give the AI real time insights into the world around you in the least invasive way possible. transparency mode allows you to hear everything as you normally do and consumers are used to having them in their ears 2. you engage with the AI and need information best shared visually, the apple watch screen provides this in a light weight way 3. the watch screen isn't enough, so and you need more consistent engagement with a screen, your phone takes you the rest of the way The goal of AI wearables is to be as minimally invasive as possible, not to push a digital overlay on top of everything you see. Further, apple has the supply chain and existing consumer relationship to dominate.
Show more
Physical AI & hardware investment is growing 20x faster than software investment this year. Building in the physical world with AI as a first class citizen is going to shape the next decade of investments. Path has been doing this for 8 years, it is not easy, it is not fast, but it is the future. The rise of embodied intelligence is happening
Show more