Joined
@EdLudlow to talk about AI progress, independent evaluation, and why I’m optimistic.
Some quick takes:
- We’re better at building AI than understanding it. Attention towards testing/evaluation matters more than slowing down.
- Our RSI Index projects models could match human researchers on the tasks we test by August 2027. Embedded evaluators can produce more accurate estimates based on internal systems.
- Public conflict masks cooperation. The labs, policymakers, and enterprises we work with want better evidence. I’ve seen enough to believe coordination is possible.
- Independence comes at a cost. We’ve rejected contracts that would compromise ours. The same group doing the testing shouldn’t also sell the solution.
- Evaluation should scale through better technology. If it becomes a bureaucratic moat for incumbent labs, we’ve failed.
- Market-based evaluation has a role with or without regulation. Competition pushes us to build better technology and keep up with the frontier.