登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

alphaXiv
@askalphaxiv
High fidelity research
参加 November 2023
84 フォロー中    53.5K ファン
"What is Missing from AI Post-Training AI" The bottleneck for autonomous AI R&D may be knowing when to abandon the current strategy, and not executing it better. So AI agents can post-train models end-to-end, but they rarely rethink the strategy they started with. Across 1,338 trajectories, agents were good at debugging, tuning, and iterating, yet almost all improvement stayed inside the initial training approach. Even 2-8x more inference compute mostly produces more optimization inside the same strategy, while human guidance can change the strategy initially but the agent eventually falls back into local refinement. So whether agents can spontaneously realize when their current strategy is wrong and decide to abandon it should be the next step forward.
もっと見る