AdityaBench findings are here!
Grok 4.7 performed the worse out of all new models from these last few days
Opus 5.5 thought hard before even outputting a decision
GPT-6 Luna actually took more time than Sol (but costed less tokens)
Unfortunately all fails! Maybe next models!