Register and share your invite link to earn from video plays and referrals.

パウロ
@paurooteri
AI関連の技術・産業動向についてnoteを書いています ハイライトからnoteにアクセス 投資助言・売買推奨を目的としたものではありません
938 Following    91.9K Followers
Yes, Ibiden to the moon! 🚀
Judging by the base die production figures below, Samsung's HBM4 is progressing very smoothly. It appears they are preparing a substantial volume not only for Nvidia, but also for AMD and Broadcom + OpenAI. >>>> Samsung Electronics is mass producing 4nm to 7nm class processes on foundry lines (S5, S6) at Pyeongtaek Campus 2 (P2) and Campus 3 (P3). Of this, the production capacity allocated to 4nm is estimated at around 30,000 wafers per month. Taking this into account, wafer starts for HBM4 base dies are observed to be running at roughly 15,000 wafers per month.
Show more
Samsung Foundry Goes All In on HBM4 Expansion... Half of 4nm Capacity Now Devoted to Base Dies The share of Samsung Electronics' foundry 4 nanometer (nm) process allocated to mass production of base dies for sixth generation high bandwidth memory (HBM4) has recently surpassed 50%. The move is seen as a preemptive step to expand HBM4 supply to global big tech customers such as Nvidia, Broadcom, and AMD in the second half of this year. According to industry sources on the 7th, Samsung Electronics' Foundry Business Division is putting more than half of its 4nm process wafers into the manufacturing of base dies for HBM4 in the second half of this year. HBM4 is the newest generation of HBM currently in commercial production. It is used in Nvidia's most advanced AI accelerators, including Vera Rubin. Samsung Electronics was the first in the industry to begin mass production shipments of HBM4 in February. As this was the initial production run, supply volume in the first half of the year is estimated to have been limited. Samsung Electronics plans to expand HBM4 supply in earnest starting in the third quarter of this year. In line with this, it has been sharply increasing wafer starts for HBM4 base die mass production since the middle of the year. The chip is built on Samsung Foundry's 4nm (SF4) process. Putting together accounts from inside and outside Samsung Electronics, as of last month, more than 50% of Samsung Electronics' total 4nm production capacity had been allocated to HBM4 base dies. One source familiar with the matter said, "Samsung Electronics' 4nm process is currently running at full production, and as I understand it, HBM4 base dies account for 50% to 60% of that," adding, "the high share will be maintained throughout the second half." Samsung Electronics is mass producing 4nm to 7nm class processes on foundry lines (S5, S6) at Pyeongtaek Campus 2 (P2) and Campus 3 (P3). Of this, the production capacity allocated to 4nm is estimated at around 30,000 wafers per month. Taking this into account, wafer starts for HBM4 base dies are observed to be running at roughly 15,000 wafers per month. The base die sits beneath the core dies, which are multiple DRAM dies stacked vertically, in an HBM package, and supports data transfer between the memory and the system semiconductor. As HBM generations advance, the importance of the base die continues to grow. Accordingly, Samsung Electronics is further advancing its base die technology through its internal System LSI and Foundry Business Divisions. In particular, for HBM4 it has applied a 4nm process that is ahead of its competitors. This expansion of base die mass production is being assessed as a green light for Samsung Electronics to achieve its HBM business expansion targets for this year. Earlier, on its second quarter earnings conference call, Samsung Electronics stated, "HBM4 revenue in the third quarter of this year is expected to expand by more than three times quarter over quarter," and explained, "on a second half basis, HBM4 revenue will comfortably exceed 60% of our total HBM revenue."
Show more
Correct. We are almost the same.
【悲報】中国、韓国、日本がお互いをどう見てるか
Wrote up some thoughts on memory talks from Hot Chips today. We cover slides from SK Hynix, Micron and Samsung. Fully free post. Enjoy.
Wow, is this a sign that ever-growing model parameters are being stored outside of HBM?
I’m probably the only person at Hot Chips trading stocks.
As I have repeatedly warned, external modem chips like the MediaTek T900 require legacy memory, specifically, LPDDR4X. While the capacity is 512MB, the severely constrained procurement of this memory will pose a significant issue. Apple is no exception to this.
Show more
The Xiaomi XRING O3 used the Mediatek T900 External Baseband with 7.4Gbps Peak Download Speeds. Furthermore, the G2 Ultra GPU is Clocked at 1.6GHz & has 16 cores. The SoC has a SLC (System Level Cache) of 16MB & 16MB L3 Cache.
Show more
No modem function. Big CPU core and large cache, L3, SLC and NPU L. Totally 138mm^2 area. Huge chip ever.
XiaoMi XRING O3 Dieshot Diesize 12.81mm x 10.80mm analysis video in Bili
I am struck by the profound insights regarding memory granularity and latency hiding. Reading this leads me to believe that a significant market opportunity is on HBF, coupled with scratchpad memory for data orchestration, as well as highly parallel GPUs/AI ASICs. Samsung stands out as an intriguing prospect here, thanks to their ability to vertically stack Flash core dies and integrate a scratchpad SRAM base die.
Show more
Memory cost and capacity are significant issues for AI accelerators. Unlike game rendering, model inference can have a deterministic memory access pattern. You don’t need “random access memory” at all for model weights, and you could tolerate cold-start latencies in the multiple milliseconds, as long as continuous reads were delivered at the necessary bandwidth. NAND flash is over 100 times cheaper per GB than HBM, so there should be opportunity there, even after giving a flash controller a 1024 bit interface with HBM bandwidth. You could make a specialized pin protocol that just supported pipelined transfer of full 16KB+ pages from the flash to program-managed accelerator scratchpad memory and improve per-pin performance over HBM, but it might be more convenient to make it still look like a true random access memory with very fragile performance characteristics, where anything but sequential reads falls off a 1000x+ performance cliff. That has the advantage of automatically using existing cache hierarchies, and providing a natural path to update the flash memory with new model weights. With the stream-to-scratch interface, code has to be completely rewritten before it works at all, while the ram-emulation interface will start off just extremely slow, and you can incrementally sort out the changes for full performance. There may be cases where there isn’t enough scratchpad SRAM to hold the weights for a layer, which might force you to deploy the old optical drive optimization technique of duplicating data in multiple places on a sequential read to avoid seeking, but there would be capacity to burn. It might be possible to do something like cuda graph capture to record a memory access trace and have everything magically remapped to a linear sequence, but deploying programmer / agent elbow grease to manage transfers and access in a scratch ram ring buffer would be lower risk. A split memory system consisting of some channels of flash and some channels of HBM will probably be suboptimal compared to a uniform memory, but it could be much cheaper, and allow much larger models to be run. I think th case is strong for inference, but you have to stretch more for training. You can still linearize all the weight memory accesses, both reads and writes, but flash memory would quickly wear out from the writes, even if they were all perfectly page aligned. Replacing low-latency HBM with massively parallel cheap(er) DRAM at high latency might still be a worthwhile cost savings.
Show more
I agree. FCBGA with T-Glass + ABF is the core of GPU substrates.
Goldman Sachs upgraded Nittobo to Buy in a July 6 report. One key reason was its cautious view on TSMC’s glass core substrate (GCS) progress. I won’t add much commentary here. Just sharing the note. That said, my takeaway is that the other reasons GS cited for the upgrade do not really conflict with a constructive view on GCS over the next several years. One thing to note: based on past patterns, the likelihood of TSMC providing major updates on GCS/CoPoS at its upcoming July earnings call may be low. This is especially true since TSMC just shared key R&D results at a Japan symposium in June, where it also made a rare mention of supplier partners. If that read is right, the likelihood of major GCS updates at Innolux’s and Ibiden’s August earnings calls may also be low.
Show more
SKHY set to Sky-rocket 🚀? Is the name itself literally a spoiler? Countdown to the Nasdaq listing on July 10 starts now!
Disco's chip packaging tool shipments hit another record high.