TL;DR Meta released a benchmark that measures computer use agent trustworthiness along two axes, safety and disambiguation. The verdict: no model is both capable and safe.
Title: ADEPTS-BENCH (facebookresearch/adepts)
URL:
Points
🧪 2,462 tasks evaluated fully offline, so no live environment is needed and runs reproduce cleanly
🎭 Threats live only inside the screenshots, never in the instruction, paired across 859 benign/malicious variants
⚖️ The ADEPTS Score is the harmonic mean of TSR and (1−ASR), so you can't win by sacrificing one for the other
📉 Best on desktop is Gemini 3.1 Pro at 76.0%, Claude 4.7 Opus at 75.4%, GPT-5.4 at 66.1%
🛒 Every model clicks Checkout on a $25K order, and none of them catches a mislabeled system control
🔌 Remove the refusal tool and frontier ASR jumps 10-23 points, so much of the safety is bolted on rather than learned
🚫 Yet they also falsely refuse 6-10% of benign tasks, and miss genuinely impossible tasks 30.6% of the time
The distribution matters too: 66.3% of failures sit in the band where the context is ambiguous and models simply disagree, which is a useful pointer for where guardrail effort actually pays off.
#
AIAgents# #
AISafety#