Hmm. Today, we ran our 2100 scored runs with Grok 4.6. Versus 4.5, the pass rate regressed from 87.3% to 85.9%, and nearly doubled the median latency (71s → 131s). Grok 4.6 also errored out more ("Something went wrong"), refused to answer benign questions (e.g., "Who in Sales has access to Google Workspace"), and produced incomplete outputs but claimed success (returning 54% of the required fields, but asserting 100%).
Could be a matter of prompt tuning, but it feels like we're at the Pareto front (at least for now) on some of these models with everyday business tasks.
Read the original post for details on the shape of data and tasks RipplingBench represents; it's very practical.