JUST IN
Grok 4.6 is the only model LatchBio tested that can refuse disguised bio threats and still do real science.
• Only system above 50% on both bars
• Refused 59.2% of disguised red-team / dangerous bio queries
• Completed 64.8% of legitimate research tasks
• Averaged 62.1% across harnesses (top three spots)
• Pathogen surveillance 53.5%, behind Opus 5, ahead of GPT-5.6 Sol
xAI: it “correctly detects and refuses dangerous queries, including maliciously obfuscated biological tasks,” without blocking beneficial science.