Anthropic is unilaterially committing to embedding evaluators to verify safety practices and report incidents. Evaluators will get access badges, company laptops, and "permissions and tools similar to those of internal employees who do comparable risk assessments."
This is a remarkable step for frontier AI auditing, transparency, safety, and security. I hope that other companies follow suit. And I look forward to continuing to build the standards, practices, tools, and policy at
@AVERIorg to make this work rigorous, effective, and universal.