With AI safety topics going mainstream, the public seems anxious and disconnected from the reality of the problem. From what I’ve seen at Vals AI, I’m optimistic we’ll coordinate toward an optimal future for AI.
Our study on RSI shows that, at their current pace, Anthropic’s models could match human researchers by August 2027. That creates urgency but gives us time to prepare.
It’s hard for the public to know whom to trust when everyone debating has their own incentives. This was the concern I had when I started Vals AI: that a multipolar paradox would emerge, where the actions of self-interested parties lead to a non-optimal outcome for the system.
To overcome this, we independently evaluate models for their real-world impact. This mirrors the role of auditing firms.
I’m optimistic because we have found rational and willing partners across the industry. Every major lab has been a great collaborator, providing us with early access to models for testing on our public benchmarks. Every member of Congress and government agency we’ve briefed has been eager to learn. Enterprises are becoming more sophisticated about adopting models based on evaluated capabilities/risks.
It hasn't been easy. We have had to earn the trust of competing groups and reject significant contracts that would have compromised our independence.
But done right, evaluation can scale with the frontier through automated infrastructure. Embedded evaluators can understand systems during development while maintaining independence. This makes it easier for new entrants to compete and promotes transparency that builds trust with the public. We’re eager to see a diverse ecosystem emerge. We’ve open-sourced our core infrastructure and published our methods, supporting peers.
This is a time of high variance. The decisions we make now will have an outsized impact on the future we arrive at. I remain optimistic we will get this right in the year ahead.