注册并分享邀请链接,可获得视频播放与邀请奖励。

elvis
@omarsar0
Founder @dair_ai • Prev: Meta AI | PhD • Learn about AI Agents for FREE here:
加入 September 2015
999 正在关注    316.6K 粉丝
Interesting paper on prompt optimization. They claim that a single-lineage prompt optimizer just matched GEPA on a smaller rollout budget. Prompt optimization has been drifting toward heavier machinery, with candidate pools, reflection trees, and Pareto-based selection. NPO keeps one lineage. At each iteration it runs the student on the current prompt, collects rollout traces and rewards, and hands a sliding window of recent iterations to a teacher model that rewrites the prompt. There is no candidate population and no search tree. On the two instruction-following benchmarks it spends 3,500 and 6,800 rollouts against GEPA's 3,593 and 6,871, and it stays broadly comparable across 22 TextArena games. The interaction with teacher strength is what makes this interesting. NPO's advantage grows as the teacher model gets stronger, which suggests optimizer-side search complexity has been compensating for weak teacher reasoning all along. Paper: Chat with Paper:
显示更多
0
9
129
12
转发到社区