pitcure also training on rubrics/tasks like Frontier Code that penalize out-of-scope changes (even good ones that fix other things). their decision to penalize totally makes sense imo, but maybe sometimes you don't want that
but it'll end up being a learned behavior in the model because it was rewarded/penalized for it
broader point is that the shape of Task/Reward during induces model behavior afterwards, and it's super hard to figure out where behavior came from across thousands of tasks & rubrics
looking at the trace + task data jointly with an agent is a good start to try reverse engineering where weird behavior could come from
^ this could be a fun bench itself if it was reliable