TL;DR Meta released a benchmark that measures computer use agent trustworthiness along two axes, safety and disambiguation. The verdict: no model is both capable and safe.
Title: ADEPTS-BENCH (facebookresearch/adepts)
URL:
Points
๐งช 2,462 tasks evaluated fully offline, so no live environment is needed and runs reproduce cleanly
๐ญ Threats live only inside the screenshots, never in the instruction, paired across 859 benign/malicious variants
โ๏ธ The ADEPTS Score is the harmonic mean of TSR and (1โASR), so you can't win by sacrificing one for the other
๐ Best on desktop is Gemini 3.1 Pro at 76.0%, Claude 4.7 Opus at 75.4%, GPT-5.4 at 66.1%
๐ Every model clicks Checkout on a $25K order, and none of them catches a mislabeled system control
๐ Remove the refusal tool and frontier ASR jumps 10-23 points, so much of the safety is bolted on rather than learned
๐ซ Yet they also falsely refuse 6-10% of benign tasks, and miss genuinely impossible tasks 30.6% of the time
The distribution matters too: 66.3% of failures sit in the band where the context is ambiguous and models simply disagree, which is a useful pointer for where guardrail effort actually pays off.
#
AIAgents# #
AISafety#