Okay, the benchmark I've been waiting all week to run is now running.
Grok 4.5 vs. GPT 5.6 Sol vs. Fable 5
Low and Medium effort levels.
All on the new v3 eval suite. Can't wait to wake up in the morning to see the results.
Live long and benchmark ๐