註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

alphaXiv
@askalphaxiv
High fidelity research. Dm for promo
加入 November 2023
72 正在關注    50.7K 粉絲
“AREX: Towards a Recursively Self-Improving Agent for Deep Research” Deep research agents often fail not because they need more search, but because they don’t know which parts of an answer are already verified and which constraints are still unresolved. This paper turns research into a recursive loop: answer, verify constraint-by-constraint, preserve evidence, then refine only the weak parts. AREX adds learned context updates and step-aware RL, achieving extremely strong results with 4B and 122B-A10B models across various agentic benchmarks.
顯示更多
0
7
212
39
轉發到社區