가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Meituan LongCat
@Meituan_LongCat
Official account of Meituan LongCat LLM 🐱 Join our Discord 👉 Subscribe our YouTube 👉
가입 August 2025
17 팔로잉 중    13.9K 팬
AI agents can now propose changes, run experiments, interpret feedback, and refine technical artifacts over many iterations. But does that make them autonomous researchers? We evaluated 7 frontier models on 36 AI R&D tasks that require sustained experimentation, covering 756 trajectories in total. Final scores only tell part of the story. We looked at how agents frame solutions, turn ideas into working implementations, retain progress, and recover from failure. We also used controlled comparisons to study experience reuse and the effect of different harness designs. Three findings stood out. 1️⃣ Strong optimization performance does not necessarily mean genuine innovation. Agents can formulate and implement practical solutions, but their strongest solutions mainly adapt or combine established techniques. Only 3 of the 252 solutions qualified as novel approaches. 2️⃣ Reliability separates current models more than peak performance. Many models can find a strong solution once. Reaching it consistently is what separates them. Similar final scores can also hide very different bottlenecks in solution framing, execution, and feedback control. 3️⃣ Experience can help or mislead, while harnesses mainly affect reliability. Accumulated experience usually improves the next solution by preserving useful discoveries, but it can also carry forward misleading conclusions or anchor agents to local optima. Harness choice mainly affects stability across repeated runs. Automated harness optimization produced gains that transferred to held out tasks and another model. Overall, current agents operate more like engineering optimizers than fully autonomous researchers. They can automate parts of the research loop, but reliable performance, effective experience reuse, and genuine novelty remain open challenges. 📄 Paper: 🌐 Project:
더 보기