and when you realise most rl envs are just LLM evals but slop volumed, a lot of things (eg misalignment) start to make a lot more sense
at some point we need to seriously have a discussion about the state of LLM evals.
reading the traces and seeing truly horrible stuff