Agentic evals are messy. A benchmark score tells you something about model performance, but it also reflects the whole system around it: the harness, sandbox, dependencies, timeouts and sometimes a loophole the agent found in the task.
That’s why trajectories matter so much to us. They show what the agent actually did and whether the score means what we think it does. We publish them to make that evidence transparent and auditable, giving the wider community more to learn from.
Watch
@aalSonOfRavi and
@ConnorBAdams go deep on all of this with
@petergostev from
@arena, including some surprisingly creative reward hacks!