Is the problem (too) advanced artificial intelligence or distorted intelligence?
Repeated security debacles from OpenAI and Anthropic are now followed by AI safety researchers leaving these companies. They claim they now (finally!) realize that the breathless race these companies are engaged in is irresponsible.
Broadly construed, the problem can be interpreted as one of lack of alignment – models are doing things that are not in line with human objectives.
But the alignment discussion often veers toward the presumption that the problem is AI models becoming too advanced, too fast.
A different interpretation is that models aren’t too advanced. Nor is there any compelling evidence that they are marching toward a superintelligence humans can’t keep up.
The problem rather may be that the way that frontier labs are training these models is leading to distorted intelligence.
The problem isn’t the capabilities (though those are clearly real), but more that the capabilities are in service of some imperfect quantitative metrics – user approval, user engagement, simple task completion metrics, various benchmark scores – over which reinforcement learning optimizes relentlessly. That this process then leads to distorted behaviors in the form of gaming the evaluation of simple completion metrics, cheating, overconfidence in wrong answers, sycophancy shouldn’t perhaps be surprising.
An analogy may help. It isn’t that we have in our hands a super car that has its own mind and wants to take control of driving because it is superior to the driver.
It is more that we have a car where the steering and the brake system don’t work. It has many of the capabilities of very good cars, and its engine, acceleration and graphic interface may be very impressive.
But if you cannot steer it properly and if you cannot hit the brakes when necessary, what could does a car do? Perhaps we shouldn’t drive it until it’s fixed.