註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

elvis
@omarsar0
Building @dair_ai • Prev: Meta AI | PaperswithCode | Elastic | PhD • Learn Harness Engineering:
加入 September 2015
1.1K 正在關注    321.3K 粉絲
Impressive paper showing how much the first retrieval step matters for deep research agents. It helps to improve GPT-5.5 from 83.1% to 90.5% on BrowseComp-Plus with the same retriever and the same agent loop. It seems that the gain comes from the opening context. The authors propose Question's Gambit which runs once, before the agent starts searching. It splits the question into clues, turns each clue into complementary searches, pools the results, and reranks them. The agent then starts its loop with that ranked set already in context. The same change lifts GPT-5.4-mini from 68.1% to 79.0% and DeepSeek-v4-pro from 71.4% to 76.9%, and roughly halves calibration error for GPT-5.5. It costs between 2.3 and 5.3 extra tool calls per question. In an error analysis, only 3 of the 79 remaining GPT-5.5 errors come from the gold document never being retrieved. The other 76 happen later, when the agent previews, opens or uses the evidence. Paper: Chat with Paper:
顯示更多
0
24
111
15
轉發到社區