没想到,有人猜Opus 5只有3T?
圈子内经常讨论的一个问题,就是 Anthropic 和 OpenAI 的模型到底有多大。之前我们也发过一些推测,今天又看到另一个推测,把几路证据加权。
他的方法大概是这样。
用Gould et al. 那套 No-CoT benchmark,不让模型写出思考过程,直接给答案,看它能默默推理到什么程度。再叠上 API 价格和 Epoch AI 的宏观算力趋势(Epoch是从芯片采购、数据中心功率、财报 capex 去估 FLOP数字)。最后调查了 20 个 AI 研究员和工程师,用来校准。
不过他这个方法待考证吧。
No-CoT reasoning 感觉很难代表模型总参数,尤其现在主要是MOE模型。
API 价格其实也很难代表什么。如果是算力flops数量的话,除了参数量,数据量也影响。
大概结果:
Claude Fable 5 4.5T
GPT-5.6 Sol 3.1T
Claude Opus 5 3.0T
Kimi K3 2.8T(已披露)
Claude Opus 4.8 2.7T
GPT-5.5 2.3T
GPT-5.6 Terra 2.1T
Claude Sonnet 5 1.6T
Grok 4.5 1.5T(已披露)
GPT-5.6 Luna 1.5T
Fable 5和Opus 5比之前大家认知的要小,之前普遍在猜测Opus5是一个5T左右的模型。
How big are the frontier models?
I tried to answer this question using a statistical model of intelligence indices, a No-CoT reasoning benchmark, API prices, compute trends, and a poll of 20 AI researchers and engineers. Here are the results: