Most of what we actually hand agents has no answer key. Write a report, book a flight, handle a refund well: nothing to parse out of a box and check against a test case.
@willccbb, a researcher at Prime Intellect, spends "Reinforcement Learning without Verifiable Rewards" on how you manufacture a training signal anyway.
@aiDotEngineer published it to YouTube.
It's a working set of techniques for turning messy production behavior into something you can train on.
- The same environment object is your eval. It also generates synthetic data for SFT, supports on-policy distillation, and works as a testbed for iterating on your harness.
- Grounding manufactures signal. Give a model source material, compare the run with it against the run without it, and that capability gap is something you can learn from.
- Production traces are the source material. You don't know the task distribution up front. A deployed agent's traces tell you what it actually is, even before you have labels.
- Work backwards from a solved state. Generate questions from documents, then throw the documents away. Break real PRs, diffs and test cases into pieces and replay them. Verify the easy problem, train on the hard one.
- Simulate the backends you can't control. For tools and web apps you can't program, a high fidelity simulator gives you full control of backend state, so you can plant the answer and get verifiability production never had.
- Judging is easier in hindsight. Look back across a finished trace, ask several models where it went wrong, and distill the agreement into rubric questions cheap enough to audit with later.
- Reward hacking comes from proxies that are undefined at the boundaries. Red team with adversarial prompt optimization, mine traces for hacks, and build up a corpus of the ones you find.
- Calibrate task difficulty deliberately. RL needs separation between rollouts, so tasks that are too easy or too hard hand you no advantage to learn from.
- Some failures only show up once you start training. Small runs on a single environment, with metrics logging how tool call patterns shift, are part of designing the environment.
Will's framing for where this goes: continual learning, where the human supplies expert judgment at the top and compute does the refining underneath.
I'm working through the published talks from AI Engineer World's Fair sharing summaries and takeaways. Follow for more!