注册并分享邀请链接,可获得视频播放与邀请奖励。

Poolside
@poolsideai
We build models for agentic coding and long-horizon tasks. Try Laguna:
加入 May 2023
2 正在关注    14.7K 粉丝
Agentic evals are messy. A benchmark score tells you something about model performance, but it also reflects the whole system around it: the harness, sandbox, dependencies, timeouts and sometimes a loophole the agent found in the task. That’s why trajectories matter so much to us. They show what the agent actually did and whether the score means what we think it does. We publish them to make that evidence transparent and auditable, giving the wider community more to learn from. Watch @aalSonOfRavi and @ConnorBAdams go deep on all of this with @petergostev from @arena, including some surprisingly creative reward hacks!
显示更多