登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Artificial Analysis
@ArtificialAnlys
Independent analysis of AI
参加 January 2024
682 フォロー中    150.8K ファン
Recent releases from OpenAI and Anthropic lead other models by a wide margin on Terminal-Bench-Science. GPT-6 Astra and Claude Opus 5.5 have a ~20-point lead over Fable 5.1 in our testing. The best-performing model we’ve evaluated so far outside of these labs is Qwen3.8 Max (0902) at 12%. Within the same model families, the most recent GPT and Claude model releases made large gains. Comparing max effort, GPT-6 Sol and Opus 5.5 both saw performance improvements on the dataset while also reducing cost per task compared to GPT-5.6 Sol and Opus 5.
もっと見る