Learnings from looking at 10,000+ agent traces in the enterprise
1/n The "jagged frontier" is here to stay
It's becoming clear that LLMs can be post-trained to solve whatever tasks we throw at them, as long as the task can be packaged into a verifiable RL environment. However, each model will be trained on a different data mix, and there is limited "cross-task transfer", leading to each "frontier" model being the best at different task distributions. We see this today with Claude/Kimi being better on visual/front-end, GPT being better on backend, and we'll likely keep seeing it because each lab will train their model on a different data mix