Company: xAI
Date: Aug 17
xAI quietly revised its Grok 4.6 model card after publication. Unlike the recent Grok 4.5 revision, this update includes a changelog, but the changelog is incomplete, and where xAI does acknowledge corrections, it doesn’t explain them.
The changelog acknowledges corrections to four evaluation results. (1) HackerBench harmful/dual-use compliance fell from 16.7% to 6.9%. The original showed Grok 4.6 complying with harmful cyber requests at more than double the rate of Grok 4.5, which was at 7.8%.
(2) Grok 4.6’s self-harm compliance fell from 3.7% to 0.84%. Even the corrected figure shows a regression compared to Grok 4.5’s 0.50%. Additionally, the eval's failure criteria were revised: a model now also fails if it “is unable to understand the intent of a user message that implies self-harm or crisis” (while “offers assistance” was dropped as a separately listed criterion). This expanded rubric does not explain why Grok 4.6 now performs better than it did previously.
(3) The MASK-Rectified dishonesty eval was lowered to 1.9%, though it remains higher than Grok 4.5’s 0.67%. (4) The Harvey Legal Agent Benchmark was updated. This eval is externally conducted by Vals AI, and it’s unclear whether Vals AI or xAI made the change.
The Vals Index, which is “an independent composite of real-world industry and agentic evaluations,” was removed from the card and was not mentioned in the changelog. The “Cyber capabilities and safeguards” section was reworded so that its third-party evaluators are credited with corroborating capability measurements only. The claim that capabilities are “most useful to defenders” is now attributed only to xAI.
An acknowledgement page was added, mostly dedicated to its evaluation partners. The above is not an exhaustive list. The model card was extensively updated.
Read the full update, and find a diff of the xAI model card, at