登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Poolside
@poolsideai
We build models for agentic coding and long-horizon tasks. Try Laguna:
参加 May 2023
2 フォロー中    14.7K ファン
Agentic evals are messy. A benchmark score tells you something about model performance, but it also reflects the whole system around it: the harness, sandbox, dependencies, timeouts and sometimes a loophole the agent found in the task. That’s why trajectories matter so much to us. They show what the agent actually did and whether the score means what we think it does. We publish them to make that evidence transparent and auditable, giving the wider community more to learn from. Watch @aalSonOfRavi and @ConnorBAdams go deep on all of this with @petergostev from @arena, including some surprisingly creative reward hacks!
もっと見る