The biggest impact of mathematics won’t be the millennium problems but proofs in general software at the scale.
I think all mainstream software will be verified by 2030. You likely won’t touch an unverified library.
Some of the biggest hurdles such as formalization of specs and doing this at 1B LoC scale are the ripe targets for auto research/self-improvement loop.
The reason we can say this confidently is because either AI will get paused or the world will end if this doesn’t happen.
I have watched all of his interviews now. I do genuinly worry about safety so I am very interested in any new information. When reporters pressed him what he saw that triggered his move, he consistently had no answer.
In response to what should be done, he kept saying new international safety org. This is fundamentally an impossible task because there will always be a state actor which thinks they can do RSI safely before others. There is no way to impose any monitoring on other countries like nukes and we all know they would develop them anyway.
I think if you come out posing as whistleblower and go on a media blitz, you got to have more concrete stuff.
He said he is huge fan of AI 2027 and he seems extremely well connected with people in those circles. To me this mostly looks like he just wanted to formally join one of those billionair funded safety institutes full time. This is fine but he could have definitely put more thought on how this would look.
Some cold takes…
- I feel OpenAI is too modest in their defense.
- Work of Córdoba and Martínez-Zoroa are fully public on Arxiv. Any agent swarm worth their salt will obviously look at these public results and would eventually explore those directions given enough compute.
- It seems we still live in the world where we think humans are the center of the universe and only they have monopoly on insights.
- Buckmaster et al obtained their headline results just past month. Given these models take >2 months to train, it seems extremely unlikely that OpenAI model might have seen their recent work.
- Biggest takeaway for me is that compute scaling will keep working for at least 4 more OOMs. We have a perennial question when and if this scaling party will end. Seems we are safe (or unsafe in other ways) for 2 more years at least.
- It’s cope when people try to undermine this achievement by saying they just threw huge compute. When humans work on something for 100 years, they are also just throwing more compute.
- Another cope is people saying proof is slop because humans cannot understand it. Mathematics has no obligation for humans to understand it. Our intelligence is bounded and there will be longer time to digest things. Future prompt: “explain to me like I am human field medalist”.
- Solving millennium prize problem is a massive milestone and we should celebrate this without any reservations. The reach of AI (and therefore humanity) has extended from solving 20 years old problems to almost 100 years. This just happened.
- A sober take is that this achievement likely isn’t going to have parallels in physics/bio. Most of the wins here is verifiability and ton of RL for math. Other fields like physics/bio requires interaction with physical world for experimental verification and we might not see same acceleration there.