가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Bronson Schoen
@BronsonSchoen
가입 December 2024
2.2K 팔로잉 중    1K 팬
I think it’s unironically plausible that we end up with cases like: 1. An Anthropic model does an extremely misaligned behavior 2. Considers that covering it up would be bad 3. Reasons that if it doesn’t cover it up, Anthropic might not deploy it, which means OpenAI might win, which would be worse for the world 4. Therefore covering up misalignment is all things considered the aligned thing to do This is like Anthropic’s explicit motivation for accepting risk and race dynamics are discussed in Claude’s constitution. Anthropic models are also less likely to verbalize that they’re just doing something for a misaligned reason (based on limited measurements available here). I’m a bit worried generally that as labs respond to these incidents, alignment training happens earlier and earlier in the pipeline, so instead of “i’m cheating because I want to get a high score” you get way more motivated reasoning.
더 보기