注册并分享邀请链接,可获得视频播放与邀请奖励。

Qiuyang Mang
@MangQiuyang
PhD student at UC Berkeley @BerkeleySky. Former ICPC World Finalists; Opinions are my own
加入 December 2025
822 正在关注    1.1K 粉丝
We integrated FrontierCS into Harbor and are releasing a preview long-horizon agent leaderboard (up to 835 turns, ~200K output tokens) with Kimi K2.6 @Kimi_Moonshot (score 46.9) and Claude Code Opus 4.7 @claudeai (43.0) 🚢. The goal: evaluate frontier coding agents in a setting where they iteratively write code, run experiments, read feedback, and improve in an extremely long loop. FrontierCS tasks are open-ended optimization problems. Each task has a continuous score. There is no single accepted output. Agents need to search for better solutions under a step/time/token budget. This makes FrontierCS a natural fit for agentic evaluation. Just plan, code, test, revise, fail, recover, and keep optimizing. Check out our blog: FrontierCS GitHub:
显示更多
0
5
133
20
转发到社区