Very well said and I think this is a very important conversation to have in public/out loud. You are thinking far too small if you look at the latest *general* model and a one-off task incredibly specific to *you/your trade* and say "oh, it's not good enough, it won't be". For years now, it's been "not good at X", or "it still does Y", and then they just continually blow past every threshold while focusing on GENERAL PURPOSE computing.
Thus far, every time we've applied serious compute and company focus on a specific domain, the models, capability, and importantly OUTPUTS/OUTCOMES get SO much better.
We are at a point where companies are simultaneously
i) undergoing one of the greatest capex buildouts of all time - building, structuring, financing, partnering
ii) each doing their own Manhattan project-style research effort to race to RSI first and fend off other competing labs at a time
iii) launch vertical strategies across MANY different domains all at once (e.g., coding, law, finance, math, etc.) across MANY form factors (e.g., assistant, terminal, agents, web, VM, etc.)
iv) model generations updating so quickly that we never even fully maximize nor saturate the capabilities of prior generations which are probably 'good enough' for 90%+ of tasks
One of those is a herculean task, let alone ALL of them at the SAME time.
We have yet to test a 100% focused frontier compute style training (reminder labs are 2-3 generations ahead) and GTM effort on a specific use case and applying as much compute as we can at saturating that (compute constrained + cost). Lone outlier perhaps is coding which we all know how that went.
Imagine what happens when they dedicate all their time and focus on what is already arguably AGI-level models at specific tasks and verticals. You think the models won't figure it out most economically viable tasks with hours of training + 1000s of employees fine tuning it specifically for that domain + compounding the multiplicative effect of clients * use case * SME they get from their customers in a said industry. Not saying we will do that, but hard to be bearish imo that they wouldn't be able to 'crack' many forms of knowledge work *if* that was their only pursuit vs. building "digital god" (and it still might happen)
More and more, thinking about what the world looks like when/where machine intelligence and depth goes far beyond what humans are capable. None of this means no work, terminator, etc. There will still be humans, companies, and plenty of work to go around, etc. But it's increasingly harder to envision a world where humans don't offload the majority of work to the machines. Intelligence was never the end all be all of human progress over time but it damn sure is an accelerant.
There are psychological and societal impacts as that flows through on what human-machine interface looks like, how people derive value (many people work = identity, or intelligence = most desirable trait), what people's idea of work looks like in the future, etc.
Time I start talking to some of those AGI philosophers, huh?
Show more
Great thread. It's all true. When I've said similar things in the past, people have accused me of hyping up what the models will eventually be capable of. That's not why I post. I don't work for any of the labs. I don't take money from any of them. I've never taken money to promote anything. I came here for one reason: to warn people about what was coming.
Everything happening now is a different tiny piece of the same pattern. You can see it everywhere if you look.
Terence Tao, almost exactly two years ago, on OpenAI's o1:
'The experience seemed roughly on par with trying to advise a mediocre, but not completely incompetent, (static simulation of a) graduate student. However, this was an improvement over previous models, whose capability was closer to an actually incompetent (static simulation of a) graduate student. It may only take one or two further iterations of improved capability (and integration with other tools, such as computer algebra packages and proof assistants) until the level of '(static simulation of a) competent graduate student' is reached, at which point I could see this tool being of significant use in research-level tasks.'
Terence Tao, four days ago:
'I mean it's it's it's amazing just how much we are willing to change everything without having any idea what's what's going to happen afterwards. It's it's extremely nonlinear dynamics. Any kind of monotone, one-dimensional thinking - well, oh, a little bit of this is good, therefore a lot of it is going to be a lot better - one of the lessons of math is that most systems don't work like that. Especially if you 10x, 100x things. So, you know, I mean, we're... we have to slow down. I mean, this is, it's insane this pace, and there's no reason to be this fast. There's no reason at all.'
He has seen it. I'm not posting this to belittle him, or what he's feeling. For I have been through it myself. I felt it four years ago, the first time I saw the shape of this. Right now the world is seeing that same shape, that same pattern, in what is happening in math. But this is not about math. OpenAI is not a mathematics company. It is an intelligence company. They didn't put serious resources into math until a few weeks ago, and look what has happened since.
This was not a matter of model capability either. It was a matter of resources, of allocation. Of compute, more of which is coming online every day. Once there is enough of it, the models will expand into more spheres of human expertise, and then into all of them. All the work of the mind. And it will play out there exactly as it is playing out now in math. Everyone will go through what Tao is going through, because all of us have something that means to us what math means to him.
But this is not about math, or art, or copyright. This is about everything, because it generalizes to everything. Pacing the Frontier is not about regulatory capture, IPOs, or crippling the competition. It's part of it, sure, but it's not the main reason. The main reason is fear. Everything happened faster over the last six months than anyone at Anthropic or OpenAI expected. If you know anyone who works there, you know this is true. This isn't a secret. The people who work there are saying it openly.
The old timelines are all blown up. RSI isn't two years out. It's not even a year away. I think we get the real thing by next summer. That's what they've seen internally, and that's the real reason for Pacing the Frontier. After we reach that, I think we will hit the next milestone really quickly. And after that, everything in this world will change. We are not ready for it. We wouldn't be ready if we had another ten years. We only make it through now with the help of extremely capable models, and I think trying to stop now would doom us all. But I have never once, in these last four years, believed that we were going to stop anyway. I don't even think we're going to slow down. We're going straight in. And the only way out is through.
Show more
gotta give credit where credit is due. openai has probably done more than any lab to drive down the cost of high quality intelligence.
open source obviously matters, but if you look at the full package from model ~quality, availability, latency, reliability, choice across intelligence levels, ease of access, & cost nobody has a better overall offering as a service.
we’ve been enjoying use these models quite a bit.
Show more
As I read this, I think they’re all well informed points of view
But think this also swings dramatically between Enterprise and Consumer
Consumers typically have no problem handing over data, don’t want an upgrade cycle forced on them from hardware, fine paying with cloud-based subs (often too many), and want the latest and greatest even if it’s beyond their needs
Enterprises want control over stack given privacy & security considerations and have more variance on the modernity of people’s tech stack
Show more
Some things I'm thinking a lot about lately:
- The search space for really capable models we'll be able to run on a wide variety of hardware will be quite large
- Sovereignty: people, businesses, countries will increasingly want to own their own stack
- People/businesses/countries will want intelligence that runs on hardware they already own
- Security security security (and privacy)
- Open source/weight intelligence increasingly proliferate
- People/businesses will increasingly care about cost savings (value maxxing over token maxxing)
- The People will value independence from any single model, inference engine, runtime, or hardware
Show more
This got me thinking as lines will eventually blur btwn what gets pushed into deterministic code as rules vs. probabilistic agents on well defined/repeatable tasks extremely fine tuned for situations with varying degrees of uncertainty, etc.
the decision tree will look interesting as organiations mature across this boundary
-->Explicit rules can move into code vs.
-->Recurring judgments on messy inputs may stay with small, specialized models vs.
--> Broader agents make sense when the steps themselves vary.
Show more
totally agree for "general purpose" models. The valid exception would be the cost comparison of a general purpose frontier model to a smaller model fine tuned for a narrow use case. The latter can still yield material cost savings, for example in the Crowdstrike data below, but only if you have a stable and sufficiently scaled use case to justify the R&D investment (not one time but ongoing to keep up w/ the frontier).
Show more