We ran DeepSeek v4 Pro 0813 on our cybersecurity benchmark, it outperformed EVERY (!) other model at finding vulnerabilities
- At pass
@3, it rediscovered 87.5% of the benchmark CVEs. Far above Opus 5 and Qwen 3.8 at 81.3%
- The tradeoff is precision. Only 65.6% of vulnerabilities reported by DeepSeek were valid. Far below GPT-5.6-Sol's 86.4%
- The model can also be unpredictable. It only finds an average of 58.3% of vulnerabilities per run. It is strongest when combining its runs' findings.
@deepseek_ai is amazing. They just outperformed every other labs with an open model
1/3 🧵