Register and share your invite link to earn from video plays and referrals.

Viv
@Vtrivedy10
applied research @LangChain Labs, prev @awscloud, phd cs @templeuniv
Joined February 2013
1.8K Following    16.3K Followers
How auditing RL Environments (ie. looking at the data) will help us explain model behavior "models are benchmark shaped" bc GRPO style RL is essentially - synthetic data generation by a model - induced by the environments + harness design - behavior (trajectories) reinforced by verifier design the "data" is RL environments & tasks, a model generates a series of tokens in its exploration of that environment trying to solve the task we don't really know what it will do, but we reward behavior that aligns with the verifier Encouraged behaviors: important: the only behavior that even has an opportunity to get rewarded is trajectories our environment "induces" or "encourages". something that isn't produced, can't be rewarded if our Harness (via a tool description) or Task design encourages lots of parallel search calls over a massive document corpus to find a few facts, the model will start copying that behavior because it got rewarded for it in the future, when some other problem is roughly shaped like that search Task, the model will mimic similar behaviors whether or not that strategy is optimal models do things that their data encouraged them to do via reward Tracing Behavior to Environments: every behavior can be fuzzily traced by the accumulation of different environment design and reward combinations the model saw over training i think we all explicitly know this but we don't have the interpretability tooling to understand induced behaviors at a fine-grained level over a massive training run this is why I'm very exicting that ppl are swarming around Trace analysis over Evals/Environments I bet a lot of "weird" or "bad" behavior comes from some quirks in environment design that inadvertantly rewarded that behavior the only thing that even has an opportunity to get rewarded Vestigial Behavior RL is in part inherently exploratory. Models need to do a bunch of stuff in rollouts to figure out what works, the verifier will reward what works But the verifier won't explicitly penalize any behavior that has a neutral effect but that co-occurs in a successful rollout this is how we get behavior that used to be helpful in some tasks, doesn't hurt current tasks, and thus persists over time similar to how humans have vestigial organs like the appendix i bet we can collectively figure out a ton about model behavior by mechanistically working backwards from the data and I'm bullish humans and agents working together will make a lot of progress on this in the next year
Show more