登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Ryan Greenblatt
@RyanGreenblatt
Chief scientist at Redwood Research (@redwood_ai), focused on technical AI safety research to reduce risks from rogue AIs
参加 September 2023
10 フォロー中    20.3K ファン
The evidence that Astra is more aligned than prior AIs seems dubious to me. Evidence appears consistent with the AI being as or more interested in score-seeking at the expense of user intent, but having beliefs+instincts that the scorer will catch a broader range of cheating behaviors. At a more basic level, the AI is extremely evaluation-aware and much less monitorable than prior AIs, making detecting misalignment much more difficult. If you train against specific reward hacks you ended up detecting, it's easy to end up papering over these reward hacks and getting an AI that is no more interested in pursuing user intent (or is barely more aligned to user intent) but which looks much better on metrics focused on detecting misbehavior. (As these metrics are from a similar distribution and the AI can just learn to pursue a somewhat different notion of score.) This concern seems live to me: based on OpenAI's discussion of their planned strategies to reduce misalignment in "The Hugging Face incident and the road ahead", it appears one of these strategies is setting up training environments where there is an apparently available reward hack (aka a honeypot) and then training against this reward hack. These sorts of methods don't seem like a robust way to ensure model behavior remains good in cases where detecting misbehavior is actually difficult (and may not be viable at all without much better oversight methods). Independent assessment of whether training is just papering over misalignment rather than actually solving the underlying misaligned motivations now seems necessary for alignment evaluations to be credible. This concern applies across frontier AI companies, not just to OpenAI.
もっと見る
Astra is our most aligned model, with substantial improvements in understanding user intent.