Register and share your invite link to earn from video plays and referrals.

The Midas Project Watchtower
@SafetyChanges
We monitor AI safety policies and web content for substantive changes. Anonymous submissions: Run by @TheMidasProj
3 Following    1.5K Followers
Company: xAI Date: Aug 17 xAI quietly revised its Grok 4.6 model card after publication. Unlike the recent Grok 4.5 revision, this update includes a changelog, but the changelog is incomplete, and where xAI does acknowledge corrections, it doesn’t explain them. The changelog acknowledges corrections to four evaluation results. (1) HackerBench harmful/dual-use compliance fell from 16.7% to 6.9%. The original showed Grok 4.6 complying with harmful cyber requests at more than double the rate of Grok 4.5, which was at 7.8%. (2) Grok 4.6’s self-harm compliance fell from 3.7% to 0.84%. Even the corrected figure shows a regression compared to Grok 4.5’s 0.50%. Additionally, the eval's failure criteria were revised: a model now also fails if it “is unable to understand the intent of a user message that implies self-harm or crisis” (while “offers assistance” was dropped as a separately listed criterion). This expanded rubric does not explain why Grok 4.6 now performs better than it did previously. (3) The MASK-Rectified dishonesty eval was lowered to 1.9%, though it remains higher than Grok 4.5’s 0.67%. (4) The Harvey Legal Agent Benchmark was updated. This eval is externally conducted by Vals AI, and it’s unclear whether Vals AI or xAI made the change. The Vals Index, which is “an independent composite of real-world industry and agentic evaluations,” was removed from the card and was not mentioned in the changelog. The “Cyber capabilities and safeguards” section was reworded so that its third-party evaluators are credited with corroborating capability measurements only. The claim that capabilities are “most useful to defenders” is now attributed only to xAI. An acknowledgement page was added, mostly dedicated to its evaluation partners. The above is not an exhaustive list. The model card was extensively updated. Read the full update, and find a diff of the xAI model card, at
Show more