登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Rohan Paul
@rohanpaul_ai
Compiling in real-time, the race towards AGI. 🗞️ Get my daily AI analysis newsletter to your email 👉
参加 June 2014
6.7K フォロー中    156K ファン
No amount of post-training cleanly fixes weak long-horizon foundations: noisy trajectories compound errors, sparse rewards misassign credit, and conflicting teachers trigger forgetting. This paper finds, ff you want agents to stay reliable over long tasks, give them clean world-model and long-trajectory training first, use OPD (on-policy distillation) when reward signals get too sparse, and avoid merging teachers with incompatible planning strategies. Suboptimal trajectories were especially damaging because small mistakes accumulated until middle and long tasks nearly collapsed. For post-training, OPD handled longer, noisier settings better than outcome-reward GRPO because teacher feedback arrived throughout the trajectory instead of only at the end. – arxiv. org/abs/2607.24720v1 Title: "The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation"
もっと見る