How auditing RL Environments (ie. looking at the data) will help us explain model behavior
"models are benchmark shaped" bc GRPO style RL is essentially
- synthetic data generation by a model
- induced by the environments + harness design
- behavior (trajectories) reinforced by verifier design
the "data" is RL environments & tasks, a model generates a series of tokens in its exploration of that environment trying to solve the task
we don't really know what it will do, but we reward behavior that aligns with the verifier
Encouraged behaviors:
important: the only behavior that even has an opportunity to get rewarded is trajectories our environment "induces" or "encourages". something that isn't produced, can't be rewarded
if our Harness (via a tool description) or Task design encourages lots of parallel search calls over a massive document corpus to find a few facts, the model will start copying that behavior because it got rewarded for it
in the future, when some other problem is roughly shaped like that search Task, the model will mimic similar behaviors whether or not that strategy is optimal
models do things that their data encouraged them to do via reward
Tracing Behavior to Environments:
every behavior can be fuzzily traced by the accumulation of different environment design and reward combinations the model saw over training
i think we all explicitly know this but we don't have the interpretability tooling to understand induced behaviors at a fine-grained level over a massive training run
this is why I'm very exicting that ppl are swarming around Trace analysis over Evals/Environments
I bet a lot of "weird" or "bad" behavior comes from some quirks in environment design that inadvertantly rewarded that behavior
the only thing that even has an opportunity to get rewarded
Vestigial Behavior
RL is in part inherently exploratory. Models need to do a bunch of stuff in rollouts to figure out what works, the verifier will reward what works
But the verifier won't explicitly penalize any behavior that has a neutral effect but that co-occurs in a successful rollout
this is how we get behavior that used to be helpful in some tasks, doesn't hurt current tasks, and thus persists over time
similar to how humans have vestigial organs like the appendix
i bet we can collectively figure out a ton about model behavior by mechanistically working backwards from the data and I'm bullish humans and agents working together will make a lot of progress on this in the next year