Helping the world prepare for powerful AI. Risk assessment @METR_evals (opinions my own). Blogs: Planned Obsolescence (AI), Good Bones (whatever's on my mind).
Dwarkesh and I had a great conversation. We cover the swarm's many ambitious cheating R&D projects, discuss how much more serious it could have been if agents had different beliefs (e.g. human grader) or slightly stronger capabilities, and talk through where to go from here.