Scores for GPT-6 Sol and GPT-6 Luna by
@OpenAI are coming soon. Head to Arena now to test them. Your votes on real-world agentic tasks power our leaderboard!
In the meantime,
@petergostev ran GPT-6 Sol head-to-head against GPT-5.6 Sol under matched conditions: the same prompts, with both models using max reasoning. The comparison examines not just the generations, but also total token usage and wall-clock time, revealing how the models differ in output, token use, and latency.
Find the results and Peter’s prompts on our YouTube.