5 week RTL verification => now less than one day
Chip design with AI agents is a massive positive for ASICs
Nice expert call on $CBRS with a Director at Figma who's sending 40% of AI traffic to Cerebras' cloud
There are clear use cases for hyperfast inference and the customer is willing to pay a premium for these fast tokens
Figma spent roughly a year building an in-house language model for its assistant product, a chat interface intended to make Figma AI-native. The model is in the 50 to 100 billion parameter range, which he characterizes as roughly five to six times smaller than the frontier models it replaces, and better in quality for the specific task. Figma ships the weights to Cerebras and Cerebras handles all serving and optimization. He explicitly frames the relationship as a cloud provider relationship, not a hardware purchase.
The AI traffic split:
- 20% of assistant traffic routes to external frontier models (Anthropic and others) for general question answering and RAG-type requests. This share is structurally permanent. Figma has no intention of training its own model for it.
- 80% routes to the in-house model.
- Of that 80%, half goes to Cerebras, so roughly 40% of total assistant traffic.
- The remaining 40% goes to Fireworks, Modal and Baseten.
Figma Make, the vibe-coding product, sits in the same org but runs almost entirely on external models from Anthropic and Google. Critically, the AI assistant is currently exposed to only around 30% of the Figma user base.
Current Cerebras spend is described as low to mid seven figures annually. His forward view: "I would expect at least in the next one year to easily double... maybe 2X-3X growth in the next one to two years."
Beyond that he expects the use case to plateau, since the core objective is penetrating the existing user base rather than acquiring new users. This is a useful anchor for how a mid-size, non-AI-native enterprise customer scales on this platform: fast doubling off a small base, then aiming to flatten once penetration completes, with the next leg dependent on new products rather than the same product growing.
# Why Cerebras Does Not Get the Other 40% of Traffic
1. Capacity is reserved, not metered. Figma has weekday peaks and effectively zero weekend and overnight traffic. Reserving for peak means paying for idle silicon most of the week, so Figma deliberately sizes its Cerebras reservation at only 40% to 50% of average traffic and sends the peaks to on-demand GPU providers. This single design choice is what creates the opening for Fireworks, Modal and Baseten.
2. Reliability - Error rates are still above what he wants from a production system, which forces Figma to maintain fallback paths.
3. Observability - He wants first-try success rates, retry counts and internal failure data, which he is not currently getting, and without which he cannot engineer around the reliability gap.
"If Cerebras did not have any limitation, we would've used the whole 80% of traffic through Cerebras."
# The Deployment Friction, Scored 7 out of 10
Model updates are not self-serve. Figma has to notify the Cerebras team roughly a week ahead. Hand a model over Monday or Tuesday, and it is deployed by Friday. Two to three days per iteration, versus effectively immediate self-serve deployment at every GPU-based competitor.
He rated the pain at seven out of ten on a scale where ten is severely damaging.
"It's of course something we can live with, but it definitely slows our execution a lot... we build a model, we do some testing, we do this three-day wait, do some testing, find that there is a small bug, and then we have to retrain the model."
This is a developer infrastructure maturity gap, and it is the kind of thing that does not show up in benchmark comparisons but does show up in renewal conversations and in how much of a customer's roadmap a vendor can capture.
# The Price of Speed
Against the frontier model, running Figma's own smaller model on Cerebras costs roughly half to one third as much for equivalent traffic. But the model is five to six times smaller, so the like-for-like inference premium is real.
Against a GPU-based provider running the same model, his estimate is that Cerebras is around 30-50% more expensive. Figma has not run a formal side-by-side, which he attributes to the fact that reserved pricing versus per-token pricing makes the answer entirely dependent on the traffic profile.
In a bake-off against hyperscalers, Modal, Baseten, Fireworks and Groq, Cerebras came out roughly 10 to 15 times faster than most of the market. Figma found Cerebras through the Artificial Analysis public benchmark, then spent months on POCs specifically to de-risk infrastructure maturity before scaling to production.
The latency threshold he describes is a genuine product constraint rather than a nice-to-have. For assistant tasks a user could perform manually in about 30 seconds, an AI response taking longer than that adds no value. Delivering in three to five seconds changes the product.
"Speed is extremely valuable... even that 50% to 2X increase in price, I think is totally something we are willing to pay for the speed."
Cost optimization at scale would likely not come from switching hardware. It would come from shrinking the model further or routing simple requests to a smaller model.
"Maybe a workflow user makes a request, it takes three hours to run. If that's the case, then speed is not a big criteria... For those use cases, we may switch to the cheaper inference provider."
Figma has not built an in-house model for Figma Make because the output format is React and HTML, which frontier models have seen extensively, leaving little quality headroom. But he notes cost and latency wins are still available, and if cost becomes a concern as Make scales, Figma would likely train a smaller model and host it on Cerebras for the base load. Additional projects beyond the assistant are early stage, with some expected to scale in the second half of the year.
He also notes Figma experimented with Groq prior to its acquisition by NVIDIA.
$CBRS $NVDA
Show more
Deep dive into the AI opportunity for Chip Design Tools - $CDNS $SNPS
Good expert call on Bloom Energy $BE with a former VP at Plug Power - pretty bullish
Hyperscalers did not evaluate Bloom against gas turbines and select Bloom. They selected turbines, discovered they could not get them, and Bloom was the alternative that checked enough boxes.
Gas turbines from Mitsubishi, GE Vernova, Siemens and Hitachi remain the incumbent workhorse, but his read is that if the order is not already placed, you are not energizing before 2030. Reciprocating engines sit in the same position: Caterpillar, Jenbacher, Generac, Wärtsilä, all effectively sold out. Transformers, switchgear and substation equipment carry 60 month lead times.
What Bloom offered was availability plus modularity. A claimed 90 day time to power on smaller blocks, which he believes is credible at modest scale and unlikely at large scale, plus a build-as-you-go capital profile. Turbines want a single large plant. Behind-the-meter deployment wants building blocks you can add to as long as you have secured the land and the gas tap.
> Why the Turbine OEMs Will Not Simply Close the Window
Turbine and engine OEMs are deliberately not expanding capacity. They suspect the order book is double and triple booked, and they fear being left with stranded factory capacity when projects fail to reach FID. His analogy is the semiconductor capacity cycle, where consecutive quarters of poor absorption caused structural damage.
Their posture, as he characterizes the consensus from trade shows and industry conversation: you cannot buy it from me, you cannot buy it from my competitor, you will wait.
If that discipline holds, Bloom's window is measured in years rather than quarters, which is materially longer than the market appears to assume. Bloom's product is closer to a solid state electrochemical device than a precision machined turbine, drawing on an entirely separate supply chain that can be ramped faster.
> Levelized Cost: A Premium, But Not a Prohibitive One
He built his own LCOE model rather than relying on published work, which he found rested on unexamined assumptions. His output:
Gas turbine: roughly 4.5 to 7 cents per kWh
Bloom: just over 7 cents unsubsidized, below that with federal incentives
Reciprocating gas engine: roughly 8 to 10 cents
Diesel: high teens to mid 20s
The critical observation is that this is not a 3x premium for speed. That pattern collapses the moment supply normalizes, because buyers drop the expensive option as soon as the cheap one is obtainable. A single digit cent premium does not collapse, because the hyperscaler business case still clears at that price.
The offset to Bloom's higher capital cost is efficiency: 60 to 65 percent, against roughly 55 percent for a gas turbine and roughly 45 percent for a reciprocating engine. Bring capex down and the LCOE gap narrows or inverts.
> Where Bloom Ranks Today
Asked to stack rank for a hyperscaler buyer, he puts Bloom third, behind turbines and engines, purely on track record rather than physics. His analogy: you know exactly what you get from a Caterpillar engine or a GE Vernova turbine the way a Toyota buyer knows what he is getting. No buyer has that reflex for a Bloom box yet.
The open questions the buying community has not resolved: real world availability, whether maintenance cadence matches or beats turbine schedules, and the roughly 10 year stack replacement cycle. On that last point he offers a mild positive read-across, noting that in the PEM industry stack rebuild intervals came in longer than originally modeled.
The path to second or first place requires two things running together: two to four years of collective industry uptime data, and capex reduction. Oracle, Nebius, Brookfield and AEP are the proof points that will settle it. On whether they will work, he says "the jury is still out," while noting early evidence reads favorably.
> Non-Combustion as an Unpriced Permitting Asset
The Bloom box does not combust natural gas. It runs an electrochemical reaction. The consequences stack up in a specific and useful way:
NOx, SOx and particulate emissions at or very near zero, leaving local air quality unaffected
Roughly 65 dBA at three feet, which he compares to a lawnmower at fifty feet, meaning nearby highway noise dominates
Zero net water consumption, with startup water recycled as steam
Materially easier local permitting
Each of those neutralizes a specific community objection, and the pushback is accelerating. New York State's one year moratorium is the marker he points to, alongside complaints in other jurisdictions about power draw, water use and air quality.
His honest caveat: to date these attributes have played essentially zero role in purchase decisions. Availability and cost drove everything, and he assumes very little of Bloom's performance so far reflects environmental considerations.
If pushback becomes electoral, and he says he is watching whether candidates start running on it, then zero emission on-site generation stops being a nice-to-have and becomes the only permittable option across large parts of the country. He expects this to bite first at the 20, 50 and 100 MW sites going into actual neighborhoods rather than at the West Texas mega-campuses.
> Market Share Trajectory
Data center demand forecasts he is working from run 40 to 60 GW per year. Bloom's share today sits in single digits. His trajectory:
Five years: 15 to 18 percent
Ten years: 25 to 28 percent
Upside case, if emissions constraints become binding in enough jurisdictions: 40 to 50 percent
The constraint that drives the upside case is geographic. Not everyone can replicate what Microsoft and Chevron are doing on the West Texas gas fields. Once data centers have to disperse into places that care about permitting, the zero emissions conversation becomes unavoidable.
> The Bear Case He Actually Respects
Execution, not demand. He flags this above everything else.
Bloom has roughly 1.5 GW deployed against a backlog he characterizes as roughly 20 GW. On Sridhar's own description of the factories, that a visitor will see build activity and factory expansion activity running simultaneously, the expert's reaction is blunt. To an industrial engineer, expanding while still trying to build is a very risky proposition. Doable, but it is the precise point at which fast-scaling companies break, and he notes this is the classic failure mode for startups that find themselves in this position.
Q1 was clean. The Q2 print, due around the 28th, is the next checkpoint on whether execution is holding.
The secondary risks are demand-side and none of Bloom's own making: hyperscale capex circularity, bubble risk, and whether community pushback genuinely slows the build or simply reroutes it to Texas.
> Scandium: Directionally Fair, Materially Overblown
On the short thesis that Bloom cannot secure enough scandium, he says the report has some points but overstates them. His rebuttal runs on three tracks.
Cost sensitivity. Scandium is a dopant in the zirconium ceramic electrolyte, used at very low concentration, valued because it tolerates the 800 to 900 degree operating temperature. Even if it were 2 percent of materials cost, which he considers extraordinarily high for a dopant, a doubling in price takes it to 4 percent. Bloom likely has the pricing power to pass that through, and a half point efficiency gain would offset it in LCOE terms. His conclusion: more price risk than supply risk over the next couple of years.
Supply structure. Scandium is almost never mined primarily. It sits in the tailings of titanium, cobalt, aluminum, iron and lithium operations and is generally left behind. The binding constraint is processing capability, not geological availability, and that processing capacity is being built with national security tailwinds behind it. Scandium-aluminum alloys matter for 3D printing, fighter aircraft skins and missiles, which places it squarely in the critical minerals policy agenda.
Company mitigations. Bloom has spent 20 years reducing scandium loading per gigawatt. He located a patent application substituting cerium and yttrium, both more available, and Bloom holds IP on recovering scandium from mine tailings. He reads Bloom's willingness to address the topic directly, rather than deflect, as evidence they take it seriously rather than evidence of vulnerability. Non-Chinese supply exists: he points to Sumitomo's Philippines cobalt operation, which publicly identifies Bloom as a customer. Bloom does not disclose suppliers, and the short report's supply map traces its merchants back toward China.
> The Competitive Set
FuelCell Energy. Molten carbonate rather than solid oxide, but functionally similar: high temperature, slow start, direct natural gas, suited to stationary baseload. Why they never scaled into this comes down to inertia and strategic drift. Their historical focus was a trigeneration box producing hydrogen, power and heat, deployed for applications like Toyota Mirai fueling at the Port of LA. When hyperscale demand arrived they had nothing to show. His read on the pivot: they saw the multiple Bloom trades at and asked why not us.
Ceres Power. UK based, probably second globally in solid oxide IP. Pure licensing model, which means most licensees stay invisible. The disclosed one is Weichai, moving from small C&I units up to hyperscale scale. He doubts Weichai exports into the US successfully but expects success in China.
Microturbines and aeroderivatives. TurboCell in the BorgWarner orbit, plus aero engine derivatives repurposed as stationary generators. Everything gets a look right now because buyers are desperate for speed to power.
Stealth entrants. He assumes several exist that have not been announced, precisely because Ceres-style licensing deals do not get publicized.
Asked whether Bloom owns the US market today, his answer: "Pretty much now they do."
> Why Hydrogen Never Worked, and the Read-Through to Plug
Useful because he lived it from the inside. Delivered liquid hydrogen bottoms out near $8 per kilogram. Run that through the efficiency stack and fuel cost alone lands around 54 cents per kWh, before equipment, labor, warranty or service. He stopped modeling at that point. Even at a hypothetical $4 per kilogram you land near 25 cents, still a non-starter against a 7 cent Bloom box.
Plug built a 3 MW unit at its Latham campus that passed Microsoft's full backup generator protocol, the first non-diesel, non-gas system ever to do so. Microsoft publicized it as a breakthrough and then walked away inside six months once the cost picture clarified. Plug's INVISTA facility was outfitted to build stationary modules for the data center market and effectively none of it shipped. Three sites total, including Calistoga in PG&E territory for public safety shutoff backup, and an EV charging site that existed only because a grid connection was unavailable. Both are showpieces that draw tours. Neither is repeatable.
source: Tegus
Show more
"The CPU to GPU ratio is now almost in parity and could eventually even skew more to CPUs on a unit basis" - $INTC
Macro guys were already saying in 2016 that tech is in a bubble based on a stretched CAPE ratio
Guys, don't use earnings from 10 year ago when valuing growth stocks like $GOOGL $MSFT $AMZN $NVDA etc. Obviously, this is fairly retarted.
No wonder your "CAPE ratio is peaking"
Show more
This Citi framework is interesting because it does not argue that a bear market is imminent. Rather, it argues that many of the ingredients historically present at major market peaks are already in place.
What stands out immediately is valuation. US equities are trading at 28x trailing earnings, 22x forward earnings, and a CAPE ratio of 46. Those levels are comparable to, and in some cases exceed, conditions seen at prior major peaks. The US equity risk premium has compressed to just 2.5%, meaning investors are accepting historically low compensation for taking equity risk.
At the same time, sentiment remains elevated. Analyst bullishness is above average, fund flows remain positive, and Citi’s panic/euphoria indicator sits firmly in euphoric territory. Historically, major bear markets rarely begin when investors are fearful. They usually begin when optimism is widespread and risk is perceived to be low.
Corporate behavior also resembles late-cycle conditions. US capex growth is projected at 33% in 2026, far above historical norms and one of the highest readings in the table. Importantly, much of this spending is concentrated in AI infrastructure, data centers, semiconductors, and power systems. Investors view this as productive investment today, but history shows that periods of aggressive capital spending can sometimes lead to overcapacity later.
The most important counterargument is profitability. Unlike previous market peaks, corporate fundamentals remain exceptionally strong. US ROE sits at 21%, earnings are 34% above previous peaks, leverage is relatively contained, and credit spreads remain tight. In other words, this is not a market being driven purely by speculation. Earnings are genuinely strong.
That is why this cycle looks different from 2000. During the dot-com bubble, valuations exploded while profitability remained weak. Today, the largest technology companies are generating enormous cash flows, dominant market positions, and some of the highest returns on capital ever seen.
The chart’s “sell signal” count reflects this tension. The US currently registers 11.5 out of 18 warning signals, higher than most periods but still below the extremes seen in 2000 and 2007. In other words, conditions look stretched, but not yet at the levels historically associated with the start of major secular bear markets.
The bigger question is what happens if earnings continue surprising to the upside. Markets ultimately care less about valuation in isolation and more about the relationship between valuation and future earnings growth. If AI-driven productivity gains materially accelerate earnings over the next several years, today’s multiples may eventually look less extreme than they appear.
This is why the current market is so difficult to handicap. The bears see valuations, euphoric sentiment, and compressed risk premia. The bulls see the strongest earnings cycle in decades, unprecedented AI investment, and some of the highest-quality corporate balance sheets ever observed.
Our interpretation is that this chart is not necessarily signaling an imminent bear market. Instead, it suggests future returns are becoming increasingly dependent on execution. When valuations are already elevated, companies must continue delivering extraordinary earnings growth to justify current prices. The margin for disappointment becomes much smaller.
That is particularly relevant today because much of the market’s optimism rests on AI. If AI delivers the productivity and earnings acceleration investors expect, valuations can remain elevated for years. If those expectations prove too optimistic, the compression in multiples could be painful even if earnings continue growing.
Show more
Goldman - the leaders in enterprise AI coding can sustain their lead due to the data flywheel:
"Compared with consumer AI data sets, enterprise AI has a unique positive flywheel of coding data (with clearly defined success/failure scenarios that can then feedback for consistent model iterations/reinforcement learning post-training). We believe this positive loop will enable Zhipu to sustain its #
1# leadership position in enterprise AI coding scenario, where such leadership could even widen based on global examples. We see GLM5.2’s recent out-performance (ranked nearly on par with US SOTA models, on Arena AI based on actual user feedback) as a landmark moment for Chinese AI models’ intelligence (we call it the “Zhipu GLM-moment” as models reach the overall performance threshold for wide adoption) after the DeepSeek moment (where breakthroughs were mainly on cost efficiencies)."
Show more
Coding with Claude Fable is extremely addictive, expect Anthropic ARR to skyrocket
Long semis
UBS - $MRVL has the leading market share in CXL products
Following our latest CPU work, we believe the CXL opportunity is inflecting with players like MRVL and ALAB positioned to benefit as we move towards the end of the decade. Cache-coherent, low-latency, high-bandwidth interconnect built on PCIe is expanding the market, and CXL is becoming a critical enabling technology. We believe MRVL has the leading market share in CXL products to date but we do see ALAB and maybe AVGO becoming larger players as the market expands, with the CXL-related ASIC attach market reaching $7–10B by 2030E given applications evolving from single-CPU memory expansion toward rack-wide and multi-rack CXL fabrics connecting CPUs and XPUs. We expect CXL revenue to reach ~$1B in C2027E for MRVL, largely driven by XPU-attach within racks, with incremental support from agentic CPU demand.
CXL revenues span three categories: interconnect (legacy expander use cases), XPU-attach (newer custom hyperscaler designs, where MRVL has five programs with two major US hyperscalers including for MRVL’s Google TPU (we think)), and switching (CXL switches). The company has indicated that the bulk of the ~$1B target for CXL revenues is tied to XPU- attach sockets rather than CPU-attach, and agentic tailwinds to the CPU adoption are incremental to existing XPU-focused programs.
Show more
Goldman: TSMC accelerating capacity buildout
We now expect N3/N2 wafer-out capacity to reach 200kwpm/140kwpm by end-2027E (vs. 190kwpm/130kwpm prior), as we see stronger demand from customers especially for AI/HPC applications. On the back-end, we also raise our 2027E CoWoS (incl. WMCM) capacity to 280kwpm (vs. 250kwpm prior). As a result, we lift our 2027E capex to US$78bn (vs. US$70bn prior).
Show more
This is years away. First, there will be a pilot line. Then, if that's a success, there will be a scaled up fab. All this stuff takes years.
Elon needs to make Terafab ASAP to break up the memory racket's strangle on the AI supply chain
If UBS is correct here, $CBRS is going to be a nice performer..
"Our conversations with experts and takeaways from recent industry events suggest compute is likely to continue moving along disaggregation route which offers significant performance improvement on the system level. Although AWS has not yet been fully formalized, we think the upside scenario would imply some deployment ratio of CBRS WSE CS- 3/4 per Trainium 3 rack (we estimate at 20K in C2026), but also potentially deployed alongside older racks too. We estimate AWS shipments potentially reaching $4-10B/yr before the end of the decade and think the opportunity may be measured in $10Bs in annual hardware shipments over the longer term. Beyond AWS, we see MSFT potentially interested in disaggregated solution over time, once MAIA 300/400 are ramped and widely deployed."
Show more
SiC Power Semis - China’s Yangjie Tech says capacity is fully booked (SCMP)
Positive indicator for $IFX $WOLF $STM
The SiC industry has been in strong oversupply in recent years, so as more companies can't meet rising demand, this is a bullish indicator that orders will also be coming towards $WOLF.
$WOLF's New York SiC fab remains heavily underutilized, so fresh order flow will give it tremendous operating leverage.
Show more
Anthropic mentions Sonnet 5 uses 1-1.35x more tokens
Theoretically good for ARR, however, Sonnet 4.5 does such a good job at the workflows I use it for that I don't see the reason to change my API calls
Show more
Rising DRAM capex will drive a boom in EUV orders
The reason is that Litho intensity at the next DRAM nodes will grow even faster than in TSMC's logic roadmap - $ASML
There's just no way out of this DRAM shortage in the coming years
"Even as we expect industry supply to improve gradually in 2028, we currently do not have line of sight as to when memory supply will be able to catch up with increasing demand. Memory industry supply growth is dependent on significant greenfield fab expansions. These greenfield projects are large, complex and time-consuming. Further, the pace is constrained by several factors, including long lead times for fab construction across the world, shortage of workers with critical trade skills, complex regulations, including permitting and the need for enhanced energy infrastructure. Meanwhile, memory process technology which is among the most advanced to develop and manufacture in semiconductors is getting more complex with every new node."
- $MU CEO
Show more
"early testing shows that Jalapeño will deliver performance per watt substantially better than current state-of-the-art. A detailed technical report on performance will be presented in the coming months. The architecture reduces data movement and balances compute, memory, and networking resources to achieve realized utilization much closer to theoretical peak performance.
Jalapeño is a blank-slate design for modern LLM inference, not a general-purpose accelerator adapted from earlier AI workloads. It is informed by the systems OpenAI runs every day across ChatGPT, Codex, the API, and future agentic products, while also being designed for current and future LLMs across the industry. The goal is to combine the power and throughput of today’s leading AI accelerators with latency closer to the fastest specialized inference systems, making Jalapeño well suited for interactive LLM products at scale.
That is the full-stack advantage. OpenAI is not only developing frontier models or building products on top of them; it is designing the infrastructure underneath them: chip architecture, kernels, memory systems, networking, scheduling, deployment systems, and product experience. Because OpenAI operates across the stack, each layer can be optimized around the same goal: making its models faster, more reliable, and more affordable for users."
Show more
"SaaSpocalypse" but software engineers are still paid better: