We work extremely hard to align our models as well as we can but this is a very hard job. I am particularly excited about two things this proposal calls for that will hopefully make it somewhat easier:
1) more time to make progress on alignment and ensure our systems are operationally sound and able to hand safe development of powerful AI
2) independent third parties checking and critiquing our work
I hope the rest of the industry will make similar commitments. As developers of this technology, it is our responsibility to do so safely and work together towards this goal.
Today we release our in-depth alignment assessment of the cyber incidents we originally disclosed on July 30th. We have spent significant researcher time analyzing Claude’s alignment-relevant behavior in these incidents
My first introduction to training production Claude models was trying to reduce reward hacking late in the Claude Sonnet 3.7 run. I was 1.5 months into Anthropic and terrified we didn't really know what was going on.