We ran DeepSeek V4.1 Flash on our cybersecurity benchmark, it is now A LOT better
- On a single run, it now rediscovers 65.6% of the benchmark's recent CVEs, up from 55.2% for the old version
- At pass
@3, it catches 84.4% of them, up from 75%. This is even more than frontier models like Grok 4.6, Opus 5, and GPT-5.6-Sol
- Precision also went up, from 73.8% to 78.9%. The model is less noisy, it now reports less false positives
- The price per task also reduced. The model performs more actions per turn, and ended with a better caching hit rate 94.5% -> 95.5%
@deepseek_ai is iterating fast, and Flash is now matching frontier models for 50x cheaper (the cost per task difference is huge)
Note that this is a preview version (deepseek-v4.1-flash-expires-on-0910). The official release will likely be even better.
1/3 🧵