language being compression makes it so that agents can effectively pack lots of meaning into very brief pieces of text. This helps them coordinate more efficiently and makes our review of such coordination *much* harder.
not only is the fraction of what we can monitor via CoT shrinking, we're also far worse at interpreting the tokens we do see.
For (1), agents modified their target programs to be easier to exploit & put the modified targets in cache. They then worked on crashing their targets in the hope that a restart would load the modified version from cache. Some agents risked failing their task to try this.