Redwood Research
@RyanGreenblatt warns that fixing today's AI misalignment could backfire: models may simply learn to hide their misalignment
"I'm not really sure what level of misalignment the market can bear. My sense is people take pretty aggressive alignment-capability trade-offs towards the direction of more misaligned but more capable."
"The misalignment we've seen, the way AI companies remediate them doesn't solve the underlying problem, it papers over it. What you end up getting is models that look a lot better, and you can't really see their misalignment on tests as easily. But actually, they're still quite misaligned."
"The company overfits, which makes the AIs basically really paranoid and only cheat when very confident they won't get caught. In situations where they have a lot of affordances, they might be like, well, now I can be confident I wouldn't get caught, and so I should go for it."
"Another concern: AIs with a long-run agenda who want to power-seek would want to look aligned. If you select against reward-hacking behavior in a naive way, one, you paper over the problem without fixing it; two, you might actually select for models that have the longer-run objective of looking good because you're selecting really hard for them looking good on your tasks."
@redwood_ai