Impressive benchmark scores from Opus 5, inching closer to Fable-level performance on
@harvey's Legal Agent Bench.
Maybe more impressively, during our early access testing we found that Opus 5 performed much better at low reasoning than prior Opus checkpoints, leading to 26% gains in token efficiency relative those prior models.
More here: