登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

alphaXiv
@askalphaxiv
High fidelity research. Dm for promo
参加 November 2023
72 フォロー中    50.7K ファン
“AREX: Towards a Recursively Self-Improving Agent for Deep Research” Deep research agents often fail not because they need more search, but because they don’t know which parts of an answer are already verified and which constraints are still unresolved. This paper turns research into a recursive loop: answer, verify constraint-by-constraint, preserve evidence, then refine only the weak parts. AREX adds learned context updates and step-aware RL, achieving extremely strong results with 4B and 122B-A10B models across various agentic benchmarks.
もっと見る