Register and share your invite link to earn from video plays and referrals.

Andrew Ho
@andrewho03
Prev: @OpenAI
194 Following    14.4K Followers
I'm actually fairly bearish on frontier lab valuations. I've never seen the reasons articulated to my satisfaction, so before I go to sleep, I wanted to quickly jot down my thinking here. The basic issue is that the labs are highly unprofitable. This may seem like a simple point, but private market valuations can be relatively irrational; however, like with $SPCX, post-IPO pricing will likely be much more punishing, especially as the standard 6-month lockup period expires and selling pressure intensifies. Many people claim that the labs have high margins. Yet even with high margins, a valuation of $1T would be justified only if the labs were doing nothing aside from serving inference (thus reducing costs only to those relevant to inference) and posting annual revenue numbers in the $100-200 billion range assuming ~80% gross margin and a 20x earnings multiple. This assumption is obviously not true, because the frontier labs have to continually spend money training the next generation of models. This is because of market competition from runner-up firms. For example, if OpenAI had paused model development last year, there would no longer be any point in paying GPT-5 API prices when you can just use Qwen or Kimi instead for much cheaper. Thus, the labs are forced to invest ever-increasing amounts of money in model training, in a way such that at any given point of time, the amount you're forced to invest in the next model is dramatically higher than the amount of money you're actually making, because even if your revenue goes up with higher model capabilities, so do your future training costs. This is a profoundly punishing dynamic which severely penalizes frontrunners. (There is also a related subpoint where frontier labs claim they can distill their leading models to win out at lower intelligence levels as well. This makes no sense because the revenue numbers involved are far too low when taking into consideration the rather low margin of such inference.) Frontier lab valuations appear largely to be based on the assumption that as you scale up, the capabilities which emerge will be sufficiently general and profound that we'll see explosive growth ( from things akin to AI agents starting and autonomously managing entire companies of subagents. But it's not clear to me that this is the case; indeed, as I mentioned in my previous post ( I believe that capabilities growth will be slower, spikier, and more data-limited than people currently assume. It may be the case that eventually we will see explosive growth of this nature with full automation of the economy, but at the very least my viewpoint implies much longer (multi-decade) timelines until we reach this point. It is not clear to me that the frontier labs will be able to operate unprofitably for so long, although I suppose maybe this foreshadows some sort of inevitable nationalization. I also want to make a broader point about technological diffusion. The reason why technological diffusion is slow isn't just because, e.g., old people take a long time to learn how to use technology (although this is of course a contributing factor to some degree). In my view, it's because when a new, revolutionary technology comes along, the ways to incorporate that technology into subsequent developments are not always obvious, and in fact they cannot necessarily be arrived at through the application of pure reason. If they could be, then perhaps frontier models, at a certain point, would have a perfect understanding of how the LLM application layer should be developed, and they would then autonomously code, deploy, and sell such a layer. But it seems more plausible to me that this diffusion is limited moreso by the hard problem of economic calculation--that is to say, the Hayekian notion through which the price system gradually promotes efficient allocation of resources and which cannot be simulated through central planning--and that even if we froze current capability levels at today's levels, it would take well over two decades to fully integrate in LLMs into our lives. Such a view is consequently rather bearish for the continued profitability of labs as it reduces their prospects for finding, say, something else comparable in profitability to coding agents, which seems to have been a somewhat lucky discovery by Anthropic to begin with. That is to say, even if you spam FDEs you aren't necessarily going to be able to just figure out the "correct" product shapes fast enough. Overall, I don't think that people have clearly reasoned through their mental models for why lab equity should be worth as much as it currently is, and that if you actually bother to write down such a model, you may not arrive at the conclusion that you want to arrive at. This isn't to say that I don't expect AI to experience a huge (industry-wide) boom in the coming decades, but just that I'm not entirely sure I would buy OpenAI or Anthropic stock at latest valuations if I were given the opportunity to do so. Of course, as an ex-lab employee, arguably this is talking against my own book; I should really be giving people more reasons to be bullish. But in the end, my influence is so small that it doesn't make a difference, so why not have some fun?
Show more
0
232
3.4K
331
Forward to community
Today is my last day at @OpenAI. I'm glad to have spent the last eight months of my life working here! I'm starting a new company focused on the production of high-quality reinforcement learning datasets: 1. The generalization ability of LLMs is clearly very poor, with "spiky" capabilities even in areas that have received tremendous amounts of investment and attention. For example, despite multiple years with tens (if not hundreds) of billions invested, even coding capabilities don't demonstrate "generality" -- even if every model can solve Codeforces questions or port C++ to Rust better than I can, I still have to manually "deslop" pull requests. 2. The vast majority of economically productive capabilities are not well represented in existing data offerings. First, there's a certain art to the design of an RL dataset which most vendors, not having upstreamed data into large training runs themselves, don't really understand. Second, and more importantly, most work is highly contextual and not easily encoded into a gradable environment; even if we can observe a "golden path" taken by a human which we believe to be good, it's challenging to understand whether alternate, counterfactual paths produce good or bad outcomes. The basic premise here is that I have a clear understanding of what labs need/want, having explicitly been on the other side and having been involved at every level from procurement all the way through training, and I'm able to provide it. I also believe that data needs will grow tremendously in the coming years, especially as frontier labs face increasing pressure toward profitability, and that they won't get the relevant capabilities "for free" through scaling alone; instead, they'll need to spend >$100B on precise, well-targeted data acquisition. Our first products will be focused on biology and statistical reasoning: 1. First, datasets that address long-horizon scientific reasoning, drawing on my work on GeneBench-Pro with @jeremyli__. Frontier models are still unable to reliably execute "messy" data analyses that require judgment, exploration, and adaptive revision (GB-Pro passrate on GPT-5.6 Sol scarcely exceeds 30%); to address this, we have the ability to generate thousands of high-quality problems with known ground truths which can be reliably graded. (In contrast, most existing RL data for bioinformatics is either massively over- or under-specified, and will probably break your model when you train on it.) Moving the "reliability gap" from 30% to >90% is obviously required for scientific acceleration, and -- despite my skepticism about generalization of RL -- is one of the *most promising datasets* conceivable when it comes to yielding generalization benefits for models' overall reasoning capabilities. 2. Second, datasets that address capabilities relevant to day-to-day workflows. Imagine a scientist snapping a picture of some experimental process or result -- say, a cell culture plate or a Western blot -- and asking Claude a question. Frontier models remain quite bad at these questions, especially those with multimodal components. But they're obviously required for acceleration of scientific discovery; before we can dream about automating science, we have to begin with shoring up these basic, generalist capabilities. Beyond these two, we hope to expand to adjacent fields (chemistry, materials science, etc.), and then even further into fields with more direct economic applicability like healthcare and white-collar office work. I strongly encourage labs with data needs to reach out. We offer industry-standard pricing and terms, and like I said -- I know how this process works, what good data looks like, and how to demonstrate to you, convincingly, that you'll be able to upstream our data into your training processes without issue. My DMs are open!
Show more
0
182
4.3K
217
Forward to community