Register and share your invite link to earn from video plays and referrals.

Macro_Lin
@LinQingV
Ex-quant & PM|AI chip design|Semis × Capital Markets|Not Financial Advice
1.5K Following    55.1K Followers
China's DeepSeek taps CITIC Securities for domestic IPO, sources say -
Never ever underestimate the power of this technology.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Show more
congrats!🎊🎉🍾🎈
I think I can officially say: we are so back
SK hynix sharply raises share of leading edge DRAM… 1c to become main process early next year SK hynix is rapidly expanding the production share of its sixth generation 10nm class (1c) DRAM. With 1c DRAM set to serve as the core die stacked into next generation high bandwidth memory (HBM4E), the company appears to be accelerating its process transition. As demand for high value added memory for servers and artificial intelligence (AI) grows, the race with Samsung Electronics and Micron to migrate to finer process nodes is also heating up. "1c process share expected to exceed 34% within the year" According to industry sources on the 7th, SK hynix's 1c DRAM share rose from around 10% in the first quarter of this year to the 13% range in the second quarter. It is projected to reach the 24% range in the third quarter and the 34% range in the fourth quarter. In the first quarter of next year, the 1c share is expected to climb to the 35% range, overtaking 1b (33% range) for the first time and becoming the company's main process. The share of 1b (fifth generation 10nm class) DRAM was found to have peaked at the 43% range in the second quarter of this year and turned downward. As the production shift to 1c gets into full swing, the share of older generation processes is shrinking in sequence. Industry estimates show that, as of the end of the second quarter, Samsung Electronics' 1c share stood in the 16% range and Micron's in the 19% range, somewhat ahead of SK hynix (13% range). However, with SK hynix stepping up the pace of its transition in the second half of this year, it is expected to overtake Samsung Electronics (31% range) on a fourth quarter basis (34% range). On its second quarter earnings conference call last month, SK hynix said that supply of DRAM built on the sixth generation 10nm class (1c) process had begun in earnest in the second quarter. The company projected that bit growth (the rate of increase in production volume) in the second half of this year would exceed the first half, driven by expanding HBM4 (sixth generation HBM) volumes and rising shipments of 1c based commodity DRAM. Process transition in preparation for HBM4E performance gains The battle for HBM4 leadership between Samsung Electronics and SK hynix is also intertwined with the pace of the 1c transition. Samsung Electronics is applying 1c DRAM from the HBM4 stage onward and is touting top tier performance with operating speeds of around 11.7Gbps. SK hynix, by contrast, chose a strategy that prioritizes mass production stability, relying on its proven 1b DRAM and advanced MR-MUF packaging technology, and is applying 1c DRAM as the core HBM die for the first time starting with next generation HBM4E (seventh generation HBM). Industry observers say this difference in strategy is being reflected in the two companies' market shares. Major research firms including Counterpoint Research project this year's HBM4 market share, on a combined basis across NVIDIA, Google, AMD and others, at the mid 50% range for SK hynix, the high 20% range for Samsung Electronics, and the high 10% range for Micron. The picture is one in which SK hynix holds its volume advantage on the strength of mass production stability, while Samsung Electronics seeks to expand share through a technical spec advantage. Against this backdrop, analysts say SK hynix's push to speed up the 1c transition will translate into tangible benefits in cost and productivity, beyond simply shrinking the node. Since bit output per wafer rises compared with the previous generation, more bits are produced from the same wafer input, improving cost competitiveness, which is cited as a factor that will help defend DRAM segment profitability from the second half onward. With the 1c transition proceeding alongside a growing mix of high value added products for servers and HBM, analysts say productivity gains are highly likely to feed directly into margin improvement. An official in the semiconductor industry said, "The pace of the 1c transition itself is encouraging, but it only becomes meaningful if actual production yields and customer qualification schedules back it up," adding, "Starting with HBM4E, the fine process competition between Samsung Electronics and SK hynix will intensify further."
Show more
It seems AI scaling is shifting from storing intelligence in ever-larger models to extracting more intelligence from fewer parameters through compute-intensive inference.
A lot of hype around OpenAI's Astra model here on my timeline today. Apparently, this goes back to a new article from The Information, which said Astra is a "recurrent depth or looped transformer". It's always interesting to read about new or different approaches (including rumors about what the closed labs may be up to), but let's debunk this a bit. About 2 months ago, I shared the architecture details of Nanbeige, for example, where "Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters." Yes, that's it. The looped transformer idea is just reusing layers in the transformer block. In the case of Nanbeige, the main idea is to reuse the same 22-layer stack (=transformer block) twice instead of once. So, effectively it extends the 22-layer architecture to 44 layers, but without duplicating the weights. In simple terms, this roughly doubles the size of the model (if we ignore the embedding and output layers for a second). But instead of requiring 2x the storage and RAM to host this model, it stays at the same size since we reuse the components. However, it's almost 2x as expensive in terms of compute, because we run the embedded text through almost 2x as many layers. Why? In the Nanbeige 4.2 technical report, the researchers found that two passes gave the best trade-off and retained about 75% of the token efficiency of a standard architecture. (More passes gave barely any gains but made the training much slower and much more expensive.) While, as far as I know, Nanbeige 4.2 is the first notable open-weight model that adopted this approach, the idea goes back to the NeurIPS paper "Mixture-of-recursions: Learning dynamic recursive depths for adaptive token-level computation". Actually, this paper proposes a mechanism that is a bit more sophisticated by adding a learned router that determines whether each token receives one, two, or more passes. So, easy tokens can exit early while harder tokens receive additional computation. In sum, Astra may be a really good model, but this shouldn't be about this "looped transformer aspect," which is just a tiny architectural tweak. Also, the statement "the new technique works in a way that obscures some or all of the AI's reasoning, otherwise known as 'chain-of-thought'" is not necessarily true with respect to the looped transformer method. It's possible that The Information journalist refers to some other technique or misunderstood the looped transformer method. Reusing layers does not by itself suppress visible chain of thought. It adds computation in hidden states before the next token is emitted, just as ordinary transformer layers do. But based on the information we have, the only plausible interpretation here is that if a model uses more of these recurrent passes, it may need to generate fewer intermediate reasoning tokens. So then more of its computation happens in latent activations that cannot be read as text. But we would get the same effect if we were scaling up the model size, like GPT 5.6 Luna -> GPT 5.6 Sol.
Show more
This is good news for ASMPT, which supplies TCB bonders to CXMT.
CHINA’S CXMT IS PREPARING TO PRODUCE ADVANCED LPDDR6 CHIPS IN SMALL QUANTITIES BY THE END OF THIS YEAR. — BBG
DigiTimes: 2027 DRAM/HBM capacity already sold out According to a DigiTimes report, annual 2027 DRAM and HBM capacity has been fully booked out ahead of schedule. NAND flash supply isn't as tight as DRAM, but bookings are expected to be completely filled by the end of this month. Per sources cited by DigiTimes, companies aren't openly acknowledging that they need to lock in memory by this month — because they worry that if more players pile into the scramble for supply, their own allocations will shrink. Manufacturers also typically deliver only 60–70% of the volumes buyers originally requested, as memory makers prioritize demand from CSPs and AI majors. As a result, the DRAM volumes smartphone and PC makers can secure in 2027 are expected to fall markedly versus 2026. DigiTimes, citing industry sources, noted that with capacity now effectively almost entirely sold out, companies that have yet to finalize their volumes could face an even more acute "can't buy it even if you want to" situation in 2027. The report added that elevated memory prices are set to become the norm. While the sharp price spikes driven by the fight for allocation may subside — meaning the magnitude of 2027 price increases could be more moderate — the overall supply shortfall and the cost burden on the finished-goods industry are unlikely to be resolved in the near term.
Show more
That’s a shame. His prime brokers had to unwind the entire book to meet margin calls. Too young, drowning in leverage.
$AMD Helios networking explained: Scale-up: inside the rack - 72 MI455X GPUs - 12 Broadcom Tomahawk 6 ASICs - 36 parallel network planes - 3.6 TB/s bidirectional per GPU - 260 TB/s aggregate The network is based on two standards: > UALoE defines how GPUs communicate, including memory reads and writes, transaction ordering, and how traffic is distributed across the 36 network planes > ESUN provides the Ethernet fabric underneath, handling forwarding, flow control, congestion management, link-level retries, and lower-overhead transmission for small messages The path is: MI455X → UALoE → ESUN Ethernet → Tomahawk 6 → MI455X The 31 TB of HBM4 stays physically distributed across 72 GPUs. UALoE lets any GPU read or write any other GPU's memory directly, making the whole rack act as a single computer Scale-out: between racks The racks are connected with Ethernet scale-out through 800G Pensando Vulcano AI NICs into a UEC-ready fabric The path is: Helios rack → Vulcano NIC → UEC Ethernet → another Helios rack
Show more
Buy dip until September ends. Mark my words.
Trump: China's Xi is coming on Sept 24th
Besi making bank in China Likely supporting LogicFolding and other 3D IC initiatives
China's CXMT is seeing strong DRAM demand, with its production reportedly reserved through 2027. 🔗 Dell, HP, Lenovo, and Apple are reportedly among the companies expected to receive priority allocations.
Show more
has completed construction of a giant data center that houses only Chinese-made chips, a big step in Beijing’s efforts to replace restricted Nvidia silicon for future AI development
Show more
This is significant. $STM is the market leader in microcontrollers and they are having problems getting allocations at TSMC for their mature node chips. Reason is that TSMC is allocating this capacity to Nvidia ecosystem and asing STM to move to 40nm of below. As a result, STM is asking its customers to input orders for whole 2027. Companies with their own fabs will benefit from this situation.
Show more
I've been discussing Huawei's τ scaling (temporal scaling) with people recently, and noticed the conversation tends to stay at the surface level without reaching its substance — likely because many participants don't come from an EE background and aren't familiar with the classical meaning of τ in circuit theory. The very first time constant you learn in a circuits course is τ = RC: the resistance of a wire multiplied by its capacitance gives the order of magnitude of the time a signal needs to traverse that wire. The longer the wire, the greater the resistance and capacitance, and the slower the signal. Within this framework, the past sixty years of geometric scaling are reinterpreted as one particular implementation of temporal scaling. Transistors were shrunk to shorten switching delay; circuits were packed more tightly to shorten metal interconnects and reduce signal propagation delay. Geometric scaling was only ever the means — compressing delay was always the end. Huawei's thesis is that once geometric scaling stalls, you find other ways to keep compressing delay. As it happens, He Tingbo's τ scaling paper released its v2 a couple of days ago, expanding from 16 to 23 pages. I compared the two versions: the data and conclusions are unchanged. The additions are essentially responses to several points of criticism the industry raised about v1. Three are worth discussing. The most important addition is the test evidence now backing the previously bare claim of "41% energy efficiency improvement." In v1, that number had no baseline and no test conditions — the most obvious target for scrutiny. V2 supplies a full comparison table. The baseline is the 2025 Kirin 9030 Pro. Both chips use the same mature process node; the key difference is that the baseline uses a conventional planar design, while Kirin 2026 folds critical paths across two vertically bonded wafers. Folding shortens interconnects and reduces interconnect delay. The timing margin freed up on the critical path translates directly into a higher maximum clock frequency: 3.1 GHz at 1.1 V supply, 13% above the baseline. The "41% energy efficiency improvement" comes from a separate operating point specifically configured for an iso-performance comparison: voltage scaled down to 0.9 V, frequency scaled down to 2.5 GHz, with measured power at 25°C coming in at 0.59× the baseline. A back-of-the-envelope estimate checks out: dynamic power scales roughly with the square of supply voltage, so an 18% voltage reduction contributes about one-third of the power drop from the square term alone. Factor in the 9% frequency reduction and the interconnect capacitance eliminated by folding, and you land right around 0.59×. So the precise meaning of "41% energy efficiency improvement" is power reduction at iso-performance. In essence, the timing margin gained from folding is traded for lower power consumption; the efficiency gain comes from logic folding. As a side note, v2 also reports that power density after dual-layer stacking is actually 5.6% lower than the baseline. The second addition addresses the question peers are most likely to ask: 3D stacking has been around for years — AMD's 3D V-Cache and Intel's Foveros are both in volume production — so what's new about LogicFolding? To understand the paper's answer, you first need to know how two layers of silicon communicate. They rely on inter-layer bond pads, which function like elevators connecting the upper and lower floors. In prior production 3D stacking, bond pad pitch ranges from 9 μm to tens of micrometers, yielding roughly ten thousand connections per square millimeter — enough to attach a bus to an entire cache block. So the established design approach has been to move complete functional blocks wholesale onto the upper tier. AMD, for example, stacks an entire cache die on top of a processor die; the two tiers are designed independently and connected through an interface. But inside a chip, a single square millimeter contains hundreds of millions of transistors. If you want adjacent logic gates to sit on different tiers — one on top, one on the bottom — that connection density falls far short. Kirin 2026 brings bond pad pitch down to 1.5 μm, yielding 440,000 connections per square millimeter. That approaches the density of the top-level metal wiring inside a chip. Routing a signal across tiers costs roughly the same as routing it across metal layers within a single die. At this point, the two silicon layers merge into a single entity in the circuit sense. EDA tools can decide at the individual logic-gate level which gate goes on which tier, handing the problem to algorithms for global optimization — a completely different degree of design freedom from what came before. The paper also explains why they didn't take the more aggressive route of fabricating a second device layer directly on top of the first. That approach offers the finest inter-layer connectivity, but manufacturing the second layer requires high temperatures that damage the already-completed first layer. It isn't production-viable today. The third addition is thermal management. Vertical stacking significantly increases thermal density per unit area, and the lower die's heat dissipation path is blocked by the upper die. This is the first objection anyone raises about 3D stacking, and v1 did not address it in depth. V2 openly acknowledges that thermal management remains a key challenge for the LogicFolding architecture. The countermeasure is thermally-aware partitioning and floorplanning: during the design phase, high-power circuits are excluded from folding candidates, and the floorplan avoids placing high-power blocks in vertical adjacency to prevent hotspot superposition. Whether this strategy is a set of manually imposed engineering constraints or has already been codified into an automated flow within their internal EDA tools, the paper does not say. It only identifies a multi-physics tool chain as the single most important investment for the next decade. Combined with the measured data showing power density 5.6% below the baseline at the iso-performance operating point, the thermal concern has at least received a direct response. That said, this approach is fundamentally avoidance-based. As stacking grows to three or four tiers, the design space eligible for folding will be progressively squeezed by thermal constraints — a boundary the paper does not explore. Additionally, v2 includes a cross-sectional micrograph of the bond interface between the two wafers and explicitly states that wafer-on-wafer hybrid bonding is used. This spec is worth benchmarking against the industry: 1.5 μm pitch wafer-to-wafer hybrid bonding on a production logic chip has no precedent. TSMC's SoIC is currently in production at 6 μm pitch; Intel's Foveros Direct is at 9 μm. Impressive, to say the least. After comparing the two versions, I'm left with two questions. One is about equipment: who supplied the bonding tools capable of this spec? The paper says only that it is the result of years of process development across a multi-vendor ecosystem. The other is about EDA: designing two wafers as a single chip is beyond what any commercially available EDA tool can do today. The paper acknowledges this, stating only that methodological details will be "published within months." Yet the frequency table shows that the 2027-generation Kirin at 3.39 GHz is already tagged as having physical silicon, meaning this toolchain was up and running inside Huawei long ago — and has been validated on at least two product generations. My personal guess is that this EDA capability was built in-house by Huawei. If anyone has insight on this, I'd welcome the discussion.
Show more
It’s just beginning of dumping cheap token if we talking about two or three years time frame.
I don’t understand why supporters of Chinese open source keep pretending not to see what is happening right now: even Chinese “open-source” players are raising prices or gradually shifting toward more closed models. I think China’s token dumping has already bottomed. Just look at Zhipu, the company behind GLM-5.2. Even Zhipu has raised prices several times this year. Chinese LLM developers cannot ignore ROI forever. Open source is not the same thing as cheap API pricing.
Show more
This is not good news for Nvidia. $NVDA
INTERESTING: Only 3 months after Rubin Ultra was announced at GTC 2026, the original 4-die Rubin Ultra has been cancelled due to manufacturing execution concerns. The new “Rubin Ultra” is half the size/~ half the real-world performance of the original Rubin Ultra. 1/4🧵
Show more
WFE = 📈 The angstrom era of semis will demand chiplets thus all or most silicon will utilize advanced packaging. Huge need for new tools. Key nuggets from $AMAT Master Class on Advanced Packaging: - AMAT is increasingly an advanced packaging story, not just a front-end WFE story. - Five of six new platforms were aimed at back-end / packaging processes, showing where the innovation roadmap is shifting. - Advanced Packaging should exceed $2B in 2026E and grow >50% Y/Y, driven by HBM, panel-level packaging, hybrid bonding, and new integration schemes. - DRAM WFE remains a major growth lever, with AMAT arguing DRAM spend should stay well over 2x NAND WFE. - DRAM complexity lifts AMAT revenue per wafer: roughly +10% from 6F² to 4F², and another +15% moving toward 3D DRAM. -NEXX gives AMAT a more complete panel-level packaging flow, including lithography, deposition, and electrochemical deposition. - AMAT is pushing into some Lam-adjacent sockets, especially ECD and PE-CVD tools tied to advanced packaging and HBM.
Show more