On Friday Anthropic came out with their Claude Code plugin which will do a sweep of codebases. I did some scans across different codebases both internal and of FOSS libraries and wanted to share my thoughts.
The pros:
1. If you have not done any codebase scans up to this point, this is the most straight forward experience I've found up to this point.
2. It does find issues! There was 1 or 2 things that were not picked up from other scans that I've done up to this point.
3. It does a very thorough job to screen out false positives.
4. You are able to pick different levels of effort that were clearly explained (picture below)
The Cons:
1. It is comically expensive, I ran a max scan that was ~$1,000, easily would have been a third to run this scan with open models. This was in line with the estimates I asked the plugin to provide prior to kicking off the analysis, so if you're curious yourself, just ask for a cost estimate before starting the scan at the level of effort you're interested in.
2. While the precision (true positive) rate was high, the recall (identifying of total confirmed issues) was quite low. The plugin is definitely steered to only showing confirmed true positives, but in a security landscape I'd rather see the larger universe of possible issues than be blinded from seeing the reality.
3. Even as a member of the Trusted Cyber Access program in Anthropic - I WAS STILL DOWNGRADED FOR CYBER CONCERNS USING THEIR PLUGIN SPECIFICALLY MADE FOR CYBER SCANS! 72% of tokens were used by Opus 4.8 after getting numerous downgrades. Opus 4.8 came out in May, using FOSS models that came out in the past month is clearly a better ROI when you consider they are cheaper and better performing.
This last one gives me pause, since Anthropic has now granted the ability for enterprises to enable Mythos for scans. As much as I'd love the chance to pony up for the scan, do I have to use their harness for the mythos scans?
The model is for sure a critical part of analysis, but the harness I'm finding can be just as important for being able to focus the model to find critical issues. If it is an on rails tool just like this plugin, I'm skeptical that it'd be worth it.
Between this plugin and OpenAI's Codex Security SDK, I'd go for OpenAI's. It is also a lengthy and costly scanner you can't control, but the level of insights I got between the two repos was a meaningful difference. For multiple large codebases, I've had the Codex Security Harness fail on occasion and never finish, but its findings were thorough and more additive to my findings when compared to Anthropic's plugin.