We're entering the token-efficiency era of AI.
GPT-5.6 Luna hits an intelligence score of ~51 at roughly $0.06 per task.
Claude Opus 5 on Low effort scores about the same at over $0.30 per task.
Same intelligence. ~5x the cost difference.
Pure intelligence is no longer the strongest moat. Cost-efficiency needs to be baked in as every model is "capable enough" now.
After deployment, we applied GPT-5.6 Sol to advance the frontier of efficiency by making itself more efficient to run.
The results:
- 20% lower serving costs from production GPU kernel improvements.
- 15%+ better token-generation efficiency from improved speculative decoding.