Alignment lead at Anthropic thinks there’s a >10% chance that AI could kill all humans but it’s okay because they’re the good guys so we should trust them
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.