Training an agent isn’t primarily a raw parameter or pre-training on a model problem anymore. It’s a trajectory problem.
You need multi-turn traces that show how an agent searches, uses tools, verifies constraints, recovers from dead ends, and eventually commits to an answer.
Good traces are surprisingly hard to get.
Listen to
@shardiban talk about the 3 big learnings from our paper (first link in thread)