There’s basically unanimous praise for Opus 5.5, but I think people *might* be overlooking what its existence implies about Anthropic’s internal models.
Anthropic almost certainly has substantially stronger internal models helping generate training environments. Think of the stronger internal model as the teacher and Opus 5.5 as the cheaper deployable student.
Now Opus 5.5 itself scores 55.8% on CoBench 2.1, while Anthropic estimates roughly 85% would be required to fully substitute for its research staff.
Opus 5.5 is now only - 30 percentage points away from Anthropic’s benchmark threshold for fully substituting its research staff.