Register and share your invite link to earn from video plays and referrals.

Clark Tang
@_clarktang
investing @ altimeter // be humble, never stop learning // no investment advice all views personal
475 Following    19.5K Followers
@demian_ai I'm all for CPUs but who in gods earth would leave an enterprise server on 24/7 100% of the time lol Probably 1-10% of your estimate
Impressive from OAI! Codex users straight inflecting up after 5.6 release
The reality is what we are seeing unfold is Nvidia speedrunning the creation of a synthetic hyperscaler. Apologies in advance to all the investors who are stuck in their priors that this will trigger. But what is a hyperscaler? Strip it down and it’s a scaled infrastructure collective of CPUs, networking, storage with a development platform on top. It fulfills two purposes. Financial: it pools and smooths the financial obligations of its users, renting infrastructure as opex instead of capex. And Operational: it builds software that makes consumption the underlying primitives simple by abstracting them away. The hyperscaler makes a healthy 35-40% operating margin by buying hardware at bulk pricing, pooling scale to get a lower cost of capital, and driving utilization of that hardware with software that shares and shards workloads across many customers. But in the age of AI, the atomic units of compute changed. Training (massive coherent clusters) and inference (agentic workloads) - require a fundamentally different configuration of resources. These new workloads require dramatically more accelerated compute, shifting the design target from multi-tenant utilization (the cloud era) to absolute workload performance (the AI era). The economics of the data center inverted. A giant, redundant fleet of Amazon Basics CPUs and storage doesn’t work when the job is synchronous training and one straggling node stalls the entire cluster. For inference, tokens per watt and time to first token dominate the economics, not how many VMs you can pack in a box. And none of it works in a world of limited power (at least in the West. Maybe in China). As Nvidia built more compute and sold it to the hyperscalers, it faced a fundamental problem. The hyperscalers had classic innovator’s dilemma - expecting 35-40%+ op margin, along with an underlying desire to commoditize Nvidia's 75% GMs with their Amazon Basics equivalent. Pay an ASIC vendor a 25% margin instead of Jensen’s 75%, then stack your own 40% on top! They owned the customer relationships too, enterprises developed on AWS, Azure, GCP and their data was captive there too. But most important of all, these companies moved at their own pace. They were not scrappy or hungry to operate at the pace Nvidia or the AI labs felt was necessary to build out compute to fulfill the demand in front of them. They would never look at retrofitting a 35MW site outside of Ashburn, Virginia! Meanwhile, a group of hungry entrepreneurs noticed the fat margins the hyperscalers earned renting what was basically stock Nvidia hardware with limited software on top, and started building businesses around it. Nvidia - skeptically at first - recognized that working with these partners would lead to faster development cycles and competitive fires and pressures for the ecosystem. Thus the neoclouds were born. The software these neoclouds co-developed with Nvidia were purpose built for the new workloads. They solved the new problems and requirements operating the new infrastructure needed. They were ready with hotswaps, they did predictive maintenance, they built new storage software that was built for training with cheaper ingress and egress fees, because their competitive drive was to win workloads, not to lock in enterprise data on their platform. And it was working - AI labs started preferring to work with them over the hyperscalers. Common complaints on the incumbents: too slow, too particular with how their clusters were built, virtualization and networking overlays that made GPU clusters underperform stock Nvidia reference designs. Neocloud bare metal was cheaper too as their teams built AI software, not a cloud data warehouse business. And they were happy to run at half the margin (~20%) that the big guys would never accept. But the hyperscalers still had one structural advantage: their balance sheets. Investment grade. Able to fund speculative capacity ahead of demand and rent it out at much higher spot rates. The neoclouds couldn’t play that game as lenders would only finance hardware that was already contracted with offtake. And more expensive if that offtake were the labs which at an earlier point were much more speculative. If only they could build ahead of demand, they could maybe earn the kind of returns Elon is achieving on Colossus. But the twist is that balance sheet edge is eroding in real time. Google just printed its first negative-FCF quarter and raised $50B equity. Microsoft is carrying $329B of leases signed but not yet commenced. Even the IG balance sheets hit the wall - more capital had to come from somewhere else. And that's how we got to where we are today. Look at what Nvidia has actually built. The operational half of a hyperscaler: DSX OS and Mission Control to run and operate GPU fleets, DSX reference designs and Omniverse digital twins as hardened playbooks for building a data center itself. Dynamo for inference serving. All the old secret sauces of the hyperscalers built specifically for new age data centers that they have led the way in architecting. Offered to any hungry, technically competent team with a serviceable site. And then the financing half: the revenue share and credit support model that smooths utilization across a distributed fleet the way multi-tenancy used to. Support the operator, release capacity to demand, share in upside, and now bring $500B of third-party capital to the table. Nvidia standardized the asset with reference designs, proved the compute was “fungible and transferable across customers and operators” and showed infrastructure investors DD unlevered yields across 7-8% hurdles. Those investors wet their beaks on early special situation financings, saw the paybacks, and understood the demand was global. That’s why Jensen spent 2025 flying around Europe, the Middle East, and Southeast Asia - these are the ground zero for new compute sites. The reality is that this didn’t happen just over the last 3 months. CoreWeave master agreement in 2023, the $6B spot reserve backstop in 2025 (to sponsor capacity for the inference clouds), the Blackrock AI Infrastructure Partnership in 2024, Brookfield’s $100B fund with Nvidia in 2025, KKR Helix with Nvidia in 2026. And now six independent financing platforms. Chess! So now the three fears by name. Circularity? Monday was the opposite with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR bringing third party capital, independently underwritten apart from one another, replacing Nvidia’s balance sheet rather than just extending it. Useful lives / underwritability of these assets? CoreWeave just disclussed A100s, 6 year old silicon contracted through 2029 and pushed 25% price increase on its fleet in July. The collateral is aging more like an aircraft than a smartphone as feared. Market share? If you don’t see that the platform of Nvidia and the fungibility of this compute is the reason why this is even possible - the skeptics themselves are making the bull argument. The complaint that these platforms keep capital tethered to Nvidia and away from other ASICs / accelerators… $500B that can only buy Nvidia reference architecture is a moat dressed up as a risk. So what were you doing when the first synthetic hyperscaler was built under your nose? :) All views expressed are my personal views. Does not reflect the views of Altimeter or Nvidia or anyone else. Full disclosure I/we may hold positions in companies mentioned. Purely for discourse and thinking - no financial advice.
Show more
0
141
2.1K
262
Forward to community
The world is going to look back at this creation of this AI infra Financing program at one of the most consequential moments in finance If you take a step back, you’ll see how impactful this will be for the entire ecosystem This is where real analysis makes alpha
Show more
Lol if this is true, this is super bullish for Anthropic, not bearish Spend <~5B run rate to train a frontier model People distilling is "25B" of run rate - both matched up Take $15B of excess to train N+1 Model Use said model to train N+2 Model Distillers still trying to catch up to N Model PROFIT Obviously am shit posting 1) amount is way too high, chinese model companies would be the first to tell you that number is way too high 2) economics matter 3) gap is wider than you think Still - incredibly important to have open source models for ecosystem to innovate on. All innovation cannot be locked within 2 companies But everyone's arguments about open source, economics, geopolitics, are all over the place. A true rorshach test... Show me the incentive and I'll tell you what said individual's argument will be :D
Show more
I think there is general confusion around how AI works, AI tokenomics, and ultimately *what is actually priced in* for the AI trade - and that some of the existing arguments are at odds with one another Firstly to clear this up - what Brad and Gavin are saying are completely in agreement, what Gavin is laying out here is the *mega bull case* as he so states in the first sentence of his tweet lol The base case we are all living with is that the labs are going to continue to generate a significant amount of revenue this year and next year. OpenAI was already the fastest growing company of all time (and still is)... but Anthropic has just grown *SO* fast that OpenAI's growth look slow by comparison The basic chain for all of this together is as follows: Power (generation, interconnect, regulation) -> DC Shell (construction, equipment, regulation) -> Semiconductors (compute, memory, interconnect, adv packaging, wafer capacity) -> Hardware (networking, storage) -> Software (data, infra, inference) -> Models (open, closed, agentic loops, harness) How each of these interact with one another affects the ultimate cost - which is model cost Consider the following: Nvidia manufactures the bleeding edge chip for training and inference. It is very good at both training, and inference. Nvidia is the largest customer of TSMC, the memory players, substrates, lasers, transceivers etc - anything you can name on. And now to soon include power into this equation. The unit of compute is fungible because the software runs ubiquitously across all clouds, multiple industries, across all models. It is bankable by increasingly more financial institutions - infrastructure PE funds, even some IG debt now - because it is ubiquitous and observable what the market is. For this Nvidia charges the highest compute margins - ~80% on hardware. Consider the labs: Anthropic and OpenAI are inferencing across a fleet of *largely Nvidia / Google TPUs w/ some incremental gains of Trainium*. There are new entrants to the field - Cerebras, AMD, and potentially some 2027 tapeouts of new ASICs - OAI Jalapeno, new start ups etc. Anthropic and OpenAI make the best models, with a dominant share of wallet $ (Assume ~$100B ARR) at an estimated gross margin of ~70%. (economic estimates vary from 40-90% depending on what you are including). But almost certainly contribution margins on model inferencing is pushing the number higher than 70%. After establishing that though, I think it's incredibly important to state that while these things seems at odds with one another, this balance is not necessarily zero sum. The thought experiment Yes it is true that if Nvidia margins were 0, OpenAI and Anthropic could offer their intelligence at cheaper rates. How much cheaper? My estimate is NVDA DC = ~12.5B / yr Amazon Basics ASIC DC = ~$6B / yr (About 1/2 the cost - so if NVDA hardware is 2x the performance, then the cost advantage goes away - and actually that ASIC is worse off bc has much worse recontracting value so arguably depreciation curve should be shorter) So really, the labs cutting NVDA out could only offer the tokens at ~50% to 60% cheaper at their own economics. Is that signficant? Certainly. Is it an OOM difference? Not necessarily - so that's why they have prudent attempts to diversify away from NVDA (it's just good business), but they continue to rely (and actually if considering Ant's share gains, are increasing their spend on NVDA - while having competing programs). In the case of Open Source vs Closed - Nvidia obviously wants the proliferation of this because by definition all OS models will run best on Nvidia hardware out of the gate. Yes NVDA hardware will be good, but they will have this lead because of everything NVDA has been doing for the last 4 years in developing their platform ecosystem from the infrastructure (partnerships, funding, neoclouds) to the software (vLLM / other inferencing sw, inference clouds, Nemotron, NIMs, Nemoclaw etc), to install base (sovereign clouds, global partnerships, neoclouds, hyperscalers, etc) - to proliferate NVDA around the world. Anywhere there is inference that exists outside of a walled garden (the proprietary labs) - Nvidia will exist. The only ones who could potentially cut NVDA out are the labs. And the value that is captured from the labs are estimated to be in the hundreds to trillions of $ - which are obviously of much value to the world if it were offered much more cheaply. Which brings us to the debate at hand -- which one is right? The truth is no one knows. You can ask the labs, you can ask Jensen - anyone who tells you definitively is just lying to you. But you can build a plausible path to the future state using a few reasoning blocks. Here's a reasoning thread (feel free to generate your own thinking): - Bull case: Spend on the world's intelligence is about $30T / yr - What would you spend to augment that, maybe worth 30-50% of that? $10-15 T as a market? - Bear case: about 30M software developers in the world each earning $100K a year = $3T spend in salary. GitHub commits up 3x = $9T of productivity on $100B of ARR? *Even if you assume 90% of this is slop and useless, you would get $900B of ROI on $100B of spend* I have more reasoning chains, but I thought this one by Jensen was compelling - but this is where we can't give too much away :) But in spirit of crowdsourcing - some other interesting ideas I have that I am still thinking about (and encourage you all to consider as well): - Optimizations always happen - the question is just to what extent and for what reason - Agentic revenues was really what unlocked step function revenue growth - if open source is really just 6mo behind, then we should see really good agentic capabilities out of open models now too - Harness and model now tightly have to be integrated - Open Source never really makes sense as a sustainable business model - businesses investing at this scale always has to find a way to monetize that - "there is no free lunch" - not just a one model fits all... the only player that has an incentive to train on the frontier and keep completely free IS Nvidia - Rev / GW of AI labs are already nearing the highest metrics ever - now to be fair Meta and GOOG never really thought of Rev / GW as metric to lead their buildouts - was always a cost to doing biz - but it's not like we are being "stupidly inefficient" with power spend now - true mkt creation - wafer constrained, power constrained world. what's the optimal move?
Show more
The mega bull case for AI infrastructure would be *if* market share shifted away from certain frontier labs with 90%+ inference margins toward cheaper models, whether open-source or closed. It would increase the ROI on AI spend for end customers by increasing intelligence per dollar, which would drive incremental token demand. Margin dollars would effectively get redistributed from the frontier labs to AI infrastructure providers. The infra winners would be those with the lowest per token cost and the winners at the model layer would be those with the highest token efficiency. There are many reasons Jensen is so focused on open source, but this is likely the most important one as I think he is probably less worried about a monopsony these days. Lower margin % at the model layer = more margin $ at the infra layer all else equal. With SpaceX and Meta being vertically integrated and possessing the #3# and #4# models respectively it is more possible than ever. Note that Grok 4.5 is ahead of Fable for some useful tasks at a much lower cost, so ranking them #3# is conservative. This is not happening yet. Cheap, mostly open source tokens are likely the majority of volume today but the majority of economic value is still accruing to the most intelligent models. Might change though. We will see.
Show more
I no longer listen to the all in pod regularly but tuning in this weekend, listening to Chamath continue to talk about no productivity from AI and the “inevitable reckoning” makes me so bullish lmao Skill issue bro
Show more
Just touched down in NYC where I will be spending the month of July 😊 If you’re interested in chatting AI, infra, semis, tokens, etc hmu! Would love to meet AI thinkers this side of the country - 🐂🧸 & anyone in between 🫡
Show more
Who do you all think has the best mgmt team between Micron, SK Hynix, & Samsung? 😊
Please can all you AI doomers stop using Claude so my Fable 5 can actually run? My requests keep getting dropped from too much traffic T_T
This is true, going open source models is way more silicon intensive, vs usage concentrating with Ant and OpenAI is more efficient, but that efficiency gets captured by the model providers. Some gets passed on to consumers, some gets absorbed by the model players This is the fundamental tension between Nvidia / Cloud hyperscalers and OAI/Ant This is the whole on prem to cloud dynamic all over again but on steroids & with much farther reaching implications
Show more
@fuckpoasting If only they realized local hosting requires way more silicon than shared infra...
We've come a long way ! Full circle moment for $SPCX Evaluating SpaceX was my first assignment at @AltimeterCap😅 Congrats to @AntonioGracias and all my former colleagues at @valor on their hard work & success in helping build one of the most consequential companies of our lifetimes And thanks to all the incredible teams at SpaceX for allowing us to partner along the way Ad Astra 🚀
Show more
$NVDA Jensen at Financial Analyst session in Taipei: "At earnings we announced an $80B share repurchase and a 25x increase in dividend and also said we would review our dividend on a regular basis. Today we plan to return of 50% or more of FCF to our shareholders this year and next year. And beyond... to infinity and beyond"
Show more
Just my my math speculation: Pre Xai deal OAI at ~2.5GW / Anthro at ~2GW - Renting NVDA ~$13B/Yr/GW (hyperscalers / CRWV financials) - Assume mix of Trainium/TPU is 50% discount per GW - Then Anthro spending ~$13B for ~2GW then = ~$1B / mo of spend So adding this deal for COLOSSUS and COLOSSUS II = $1.25B / mo of spend That means after this deal, $NVDA penetration of Anthropic compute stack = >50% instantly... That's something Jensen mentioned on the earnings call: "And so we're growing - and our coverage of Anthropic has been largely 0 until just recently. And so we're gaining share tremendously in inference" Based on studying the industry I'm pretty sure that math is right... Even assuming Trainium / TPU is 25% discount then this deal puts Nvidia compute penetration at Anthropic at 40%! anyone else have different analysis?
Show more
Anthropic is paying $1.25B a month to SpaceX for compute
Nick’s one of the best, glad he’s finally showing his smarts to the world :) @coatuemgmt
Wow, I missed this post by Gavin. Exactly on point, and excellent. Not biased, some of the points he describes are good for NVDA, bad for NVDA, good for ASICs, bad for ASICs, good for China, bad for China. But word for word, ‘insight’ encoding super high in this post. Truth dense
Show more
Much of Dwarkesh's argument hinges on this statment which *was* accurate but will be increasingly inaccurate on a go forward basis imo:    “American labs port across accelerators constantly. Anthropic's models are run on GPUs, they're run on Trainium, they're run on TPUs. There are so many things you can do, from distilling to a model that's well fit for your chips.”   As system level architectures diverge (torus vs. switched scale-up topologies, memory hierarchies, networking primitives), true portability is eroding. The Mi300 and Mi325 had roughly the same scale-up domain size as Hopper while Blackwell’s scale-up domain is 9x larger than the Mi355 scale-up domain, etc. Many frontier models are now being explicitly co-designed for inference on specific hardware like GB300 racks. Codex on Cerebras is another example. Those models run less efficiently on other systems and the performance differentials will only widen. A model that runs well on Google’s torus topology will run less efficiently on Nvidia’s switched scale-up topology and vice versa - the data traffic is fundamentally different as a byproduct of the models being parallelized across the different topologies. Google’s internal teams - and increasingly the Anthropic teams as they become the most important customer of almost every cloud - have the luxury of operating across the stack (models, chips, networking) - but that is not the case for the rest of the market and other prospective users. Anthropic is the exception, not the rule. To wit, Anthropic and Google allegedly have a mutual understanding where Anthropic can hire the TPU engineers they need every year to ensure that they can continue to get the most out of the TPU. Given the overwhelming importance of cost per token to the economics of the labs, models will be run where they run best. Most extremely large MoE models will run best on GB300s given the importance of having a switched scale-up network like NVLink for MoE inference. When training was the dominant cost for labs and power was broadly available, labs were optimizing to minimize capex dollars. Model portability was a way to create leverage over suppliers. I think that drove a lot of the focus on portability. Today, inference costs as measured by tokens per watt per dollar are everything. Inference is way more important than training costs (inference is effectively now part of training via RL). Labs are therefore now optimizing for inference. This means increasing co-design and higher go-forward switching costs for individual models between systems. I do think this explains why Anthropic and Nvidia came together: Anthropic needed Blackwells and Rubins to inference at least *some* of their models economically. And Mythos might just end up being released coincident with the availability of Rubins for inference. TLDR: as labs shift their focus from training to inference, the costs of portability and the upside of co-design to maximize tokens per watt per dollar both rise. Portability is likely to begin decreasing as a result.   I think what I might have respectfully added to Jensen’s answer is that systems evolve under local selective pressures. The evolutionary pressure in America is a shortage of watts so it makes sense for Nvidia to optimize, as an American company, for power efficiency and tokens per watt and stay on copper as long as possible. China has a surfeit of watts. Chinese AI systems are already taking advantage of this with the Huawei Cloudmatrix 384 and Atlas SuperPoD having an optical scale-up domain that is much larger than anything offered by Nvidia today at the cost of *much* higher power consumption and much lower tokens per watt. The networking primitives for this Huawei system are very different than those for Nvidia’s systems and a model that runs well on Nvidia will not run well on that system and vice versa. This means that if a Chinese ecosystem gets momentum, Chinese models might stop running well on American hardware. And when Chinese models run best on American hardware, America is in a better position as this gives America a degree of leverage and control over Chinese AI that it risks losing to an all-Chinese alternative ecosystem.   This architectural fork makes porting and distillation less effective and strengthens the pro-American national security case for selling China deprecated GPUs imo. Also I will attest that I did not wake up a loser this morning.
Show more
Doing some prep work today and this new OpenAI image model is super cracked... It's just super steerable, incredibly smart, the one-shot end product that comes out is just magical...
Show more
we will start benchmarking companies by their revenue generated per token ($/token) in the same way industrial companies look at $/kwh. not saying that other costs (labor etc) won't exist, they will just be a much smaller fraction.
Show more