opus 5 is a sign that the obsession with ālong-horizon agentsā in model training is finally backfiring
i donāt like long-horizon agents, and iāll explain why they fundamentally donāt work
some people will immediately jump out and say āskill issueā. well, show me one profitable business you built with a long-horizon agent working all by itself - iād love to learn
so far, the only thing they were able to build thatās even interesting enough for people to talk about are those 3d games that are a partial clone of something that already existed
the reason an agent was able to build a working prototype of complex games like call of duty was that a team of humans already figured out all the requirements years ago for how such games should work, what kind of controls are intuitive, what mechanics are fun etc
all those requirements were already absorbed into the model weights, so when you say ābuild me call of dutyā the model already knows the details. its long horizon execution capability can get all the requirements implemented, which i must say is indeed impressive
but now you can see - the value of long horizon execution has a prerequisite of a massive amount of high quality requirements clearly defined upfront. it took a big team of very talented humans months of effort and many iterations to define that for call of duty
now imagine games like call of duty donāt exist yet, how would we use agents to build it for the first time? we canāt say ābuild me call of dutyā any more. and thereās no way we can define months-worth of game design details upfront
weāll have to build a tiny prototype of the most basic mechanics, play with it, see if itās fun, then iterate and expand the complexity. even with the smartest humans, thatās how we work towards something great
we donāt need agents to go dark for a long time, spend tens of thousands of dollars worth of tokens, and come back with a product the agent randomly decided to build - try build something truly novel with this and youāll see it canāt come up with anything thatās actually profitable (iāll show you why in a bit)
we need a tight feedback loop where we can collaborate with the agent, plan with it, understand what itās done, question its approach, apply our judgement, give it real world feedback and iteratively arrive at a good outcome
and thatās exactly what opus 5 absolutely suck at. why? i explained it in more depth with my previous post on how RLVR works - RLVR trains the model to generate code that can pass predefined tests in an isolated environment, which is fundamentally incompatible with the idea of having human in the loop. the more we train the models with RLVR to be ālong-horizonā, the less they care about talking to humans
ok now - why do they have to talk to humans? why canāt the models iterate and apply judgement by itself?
maybe one day they could, but not today, due to many limitations. two examples -
1. LLMs today canāt āwatch a videoā yet. they can look through a lot of screenshots, which is extremely inefficient at observing a high fps animated signal. so anything that requires continuous visual attention is something LLMs canāt do very well
2. LLMs donāt truly understand whatās āintuitiveā or āpleasantā for humans. they know whatās already proven to be intuitive and pleasant in the past, but if you present a truly novel concept, it canāt predict whether humans will like it accurately
because of those limitations, human judgment is still needed for almost anything valuable. without humans in the loop, agents will only be able to repeat something that already existed, or go in random directions without true understanding of whether itās building something useful
in summary, long horizon agents assume requirements all exist upfront. they are fundamentally against human in the loop. and they donāt have true judgement for what humans like
that, my friend, is why they donāt work
Show more