Everyone benchmarks intelligence. Nobody benchmarks honesty.
Grok 4.6 has the lowest hallucination rate of any frontier model right now. GPT-5.6 Sol sits at 92%.
The smartest model in the room means nothing if it cancels your Stripe subscriptions during a migration.