註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

alphaXiv
@askalphaxiv
High fidelity research
加入 November 2023
84 正在關注    53.5K 粉絲
"What is Missing from AI Post-Training AI" The bottleneck for autonomous AI R&D may be knowing when to abandon the current strategy, and not executing it better. So AI agents can post-train models end-to-end, but they rarely rethink the strategy they started with. Across 1,338 trajectories, agents were good at debugging, tuning, and iterating, yet almost all improvement stayed inside the initial training approach. Even 2-8x more inference compute mostly produces more optimization inside the same strategy, while human guidance can change the strategy initially but the agent eventually falls back into local refinement. So whether agents can spontaneously realize when their current strategy is wrong and decide to abandon it should be the next step forward.
顯示更多
0
6
124
11
轉發到社區