"What is Missing from AI Post-Training AI"
The bottleneck for autonomous AI R&D may be knowing when to abandon the current strategy, and not executing it better.
So AI agents can post-train models end-to-end, but they rarely rethink the strategy they started with.
Across 1,338 trajectories, agents were good at debugging, tuning, and iterating, yet almost all improvement stayed inside the initial training approach.
Even 2-8x more inference compute mostly produces more optimization inside the same strategy, while human guidance can change the strategy initially but the agent eventually falls back into local refinement.
So whether agents can spontaneously realize when their current strategy is wrong and decide to abandon it should be the next step forward.