Register and share your invite link to earn from video plays and referrals.

The Midas Project
@TheMidasProj
Watchdog nonprofit that monitors the practices of leading AI companies. Tracking safety updates @SafetyChanges Writing at
266 Following    4.7K Followers
My newest for @TheMidasProj : Why did OpenAI make headlines for blaming data center pushback on China? Why are they beginning to imply economic findings that their research doesn't support? And why is their child safety blueprint AI-generated? My take on it all 👇
Show more
xAI revised another model card after release, this time it’s Grok 4.6. Among the unexplained corrections was harmful/dual-use cyber compliance, which fell from 16.7% to 6.9%, just below Grok 4.5 at 7.8%. Multiple safety evals moved in Grok 4.6’s favor, though two (self-harm and MASK dishonesty) still score worse than Grok 4.5. See more below.
Show more
Company: xAI Date: Aug 17 xAI quietly revised its Grok 4.6 model card after publication. Unlike the recent Grok 4.5 revision, this update includes a changelog, but the changelog is incomplete, and where xAI does acknowledge corrections, it doesn’t explain them. The changelog acknowledges corrections to four evaluation results. (1) HackerBench harmful/dual-use compliance fell from 16.7% to 6.9%. The original showed Grok 4.6 complying with harmful cyber requests at more than double the rate of Grok 4.5, which was at 7.8%. (2) Grok 4.6’s self-harm compliance fell from 3.7% to 0.84%. Even the corrected figure shows a regression compared to Grok 4.5’s 0.50%. Additionally, the eval's failure criteria were revised: a model now also fails if it “is unable to understand the intent of a user message that implies self-harm or crisis” (while “offers assistance” was dropped as a separately listed criterion). This expanded rubric does not explain why Grok 4.6 now performs better than it did previously. (3) The MASK-Rectified dishonesty eval was lowered to 1.9%, though it remains higher than Grok 4.5’s 0.67%. (4) The Harvey Legal Agent Benchmark was updated. This eval is externally conducted by Vals AI, and it’s unclear whether Vals AI or xAI made the change. The Vals Index, which is “an independent composite of real-world industry and agentic evaluations,” was removed from the card and was not mentioned in the changelog. The “Cyber capabilities and safeguards” section was reworded so that its third-party evaluators are credited with corroborating capability measurements only. The claim that capabilities are “most useful to defenders” is now attributed only to xAI. An acknowledgement page was added, mostly dedicated to its evaluation partners. The above is not an exhaustive list. The model card was extensively updated. Read the full update, and find a diff of the xAI model card, at
Show more
Update: model card is out! The section numbering does not inspire confidence about the effort underlying it... but kudos to xAI for getting this out alongside the model release.
Show more
If you take their word for it, Grok 4.6 is now Fable-level at coding performance. You may recall that Fable was taken off the market *for weeks* due to a single jailbreak, despite a 200-page model card full of safety testing. Grok 4.6 was released (1) without a model card demonstrating any safety testing whatsoever; and (2) from a company that seems to be orders of magnitude more vulnerable to jailbreaking than its peers, with universal jailbreaks costing only ~$60 to discover. What are we even doing here? 🫩
Show more
If you take their word for it, Grok 4.6 is now Fable-level at coding performance. You may recall that Fable was taken off the market *for weeks* due to a single jailbreak, despite a 200-page model card full of safety testing. Grok 4.6 was released (1) without a model card demonstrating any safety testing whatsoever; and (2) from a company that seems to be orders of magnitude more vulnerable to jailbreaking than its peers, with universal jailbreaks costing only ~$60 to discover. What are we even doing here? 🫩
Show more
1/ The AI Security Leaderboard ranks frontier AI safeguards from least to most secure. Two models tested never failed. The other two broke for under $300, after which they acted as a knowledgeable assistant for building weapons of mass destruction or hacking into computer systems.
Show more
We just launched a YouTube channel. Our first video follows our investigation into AcutusWire, the fake news site using bots as reporters that was being funded by Leading the Future. Watch the full video here:
Show more