註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

pilvar (Philippe Dourassov)
@pilvar222
AI Pentest Lead @AikidoSecurity
加入 October 2013
592 正在關注    3.9K 粉絲
We ran Qwen3.8-Max on our cybersecurity benchmark. Given enough attempts, it found more CVEs than most frontier models, tying Opus 5 for first place. Here are the results: - Across the three runs, Qwen found 26 of 32 CVEs, reaching 81.25% pass@3 recall, outperforming GPT-5.6-Sol and matching Opus 5 at a lower price - Only 10 CVEs were found in all three passes, compared with 19 for both Opus and Sol. Qwen is very inconsistent, but can be very strong if run multiple times - The full campaign cost $821.35. That is 2x cheaper than Opus and Sol, but 5x more expensive than DeepSeek V4 Flash Our benchmark tests the models on finding recent CVEs from last month, so they are absent from the training data. Yet, Qwen often tried to recall historical CVEs based on the release version from source, before shifting to code-based reasoning. 1/3
顯示更多
0
25
453
44
轉發到社區