SOME DETAILS FROM OPUS 5.5 BLOG:
> most cyber tasks still get routed to Opus 4.8
> if often knows it’s being evaluated
> that makes real world behavior harder to assess
> best alignment results yet on 2,000 scenarios
> 85% fewer attempts to cross containment boundaries
> matches or beats Mythos 5.1 on biology
> they wants to rely less on reading CoT and more on interpretability
> tightening RL env filtering as a major source of misalignment
> they still argue for government regulations
> text watermarking is included