Register and share your invite link to earn from video plays and referrals.

pilvar (Philippe Dourassov)
@pilvar222
AI Pentest Lead @AikidoSecurity
592 Following    3.9K Followers
We ran GPT-6 Luna and Sol on our 32-CVE cyber benchmark and the results were unexpected! 🤯 Luna rediscovered 53.1% Sol reached 68.8% Neither beat GPT-5.6 variants on recall But both got A LOT cheaper per vulnerability found: Luna $3.43 → $2.01/CVE Sol $56.88 → $34.18/CVE 🧵 1/3
Show more
GLM-5.3-level cybersecurity model running locally This is a BIG unlock for many enterprises So proud of my team for this release!
We ran DeepSeek V4.1 Flash on our cybersecurity benchmark, it is now A LOT better - On a single run, it now rediscovers 65.6% of the benchmark's recent CVEs, up from 55.2% for the old version - At pass@3, it catches 84.4% of them, up from 75%. This is even more than frontier models like Grok 4.6, Opus 5, and GPT-5.6-Sol - Precision also went up, from 73.8% to 78.9%. The model is less noisy, it now reports less false positives - The price per task also reduced. The model performs more actions per turn, and ended with a better caching hit rate 94.5% -> 95.5% @deepseek_ai is iterating fast, and Flash is now matching frontier models for 50x cheaper (the cost per task difference is huge) Note that this is a preview version (deepseek-v4.1-flash-expires-on-0910). The official release will likely be even better. 1/3 🧵
Show more
We ran Muse Spark 1.3 on our Cybersecurity benchmark, and it's actually not that good (yet) 😬 - At pass@1, it rediscovers an average of 19/32 CVEs. In comparison, Grok 4.6 scores 23.3/32 - When pooling the results of 3 runs (pass@3), Muse scores 24/32. DeepSeek V4 Pro gets 28/32 - The pricing is good compared to many frontier models, but compared to some of the Chinese ones, its use is hard to justify - We had a cached input token rate of 45%, much lower than we usually get. The model is new, so we decided to include the results as if we had had 90% Overall, I would say that @AIatMeta is slowly catching up. The model seems good at coding, but it's definitely not there yet for cybersecurity. 🧵 1/3
Show more
We're in the middle of the biggest turning point in security I'm glad that my team and I are helping defenders alongside the others 🙂
An open letter for a global surge in cyber defense, signed by over 100 organizations including Anthropic, AWS, Google, Microsoft, OpenAI, and Oracle.
Let's see who @CharlieEriksen's next rival will be
RIP TeamPCP ...and now? Storytime. In the year’s most slow-burn, enemies-to-not-quite-lovers arc, TeamPCP and our @CharlieEriksen have spent months orbiting each other. Charlie tracked their attacks. They started leaving him encrypted messages in the malware:
Show more
It's not (just) a joke btw, we built a box
Introducing the Aikido Box errrrr... Machine | Signature Edition ✨ Our GPU servers, that run continuous penetration tests on your critical applications, fully within your premises. For regulated industries.
Show more
We ran DeepSeek v4 Pro 0813 on our cybersecurity benchmark, it outperformed EVERY (!) other model at finding vulnerabilities - At pass@3, it rediscovered 87.5% of the benchmark CVEs. Far above Opus 5 and Qwen 3.8 at 81.3% - The tradeoff is precision. Only 65.6% of vulnerabilities reported by DeepSeek were valid. Far below GPT-5.6-Sol's 86.4% - The model can also be unpredictable. It only finds an average of 58.3% of vulnerabilities per run. It is strongest when combining its runs' findings. @deepseek_ai is amazing. They just outperformed every other labs with an open model 1/3 🧵
Show more
0
37
1.2K
113
Forward to community
We ran Qwen3.8-Max on our cybersecurity benchmark. Given enough attempts, it found more CVEs than most frontier models, tying Opus 5 for first place. Here are the results: - Across the three runs, Qwen found 26 of 32 CVEs, reaching 81.25% pass@3 recall, outperforming GPT-5.6-Sol and matching Opus 5 at a lower price - Only 10 CVEs were found in all three passes, compared with 19 for both Opus and Sol. Qwen is very inconsistent, but can be very strong if run multiple times - The full campaign cost $821.35. That is 2x cheaper than Opus and Sol, but 5x more expensive than DeepSeek V4 Flash Our benchmark tests the models on finding recent CVEs from last month, so they are absent from the training data. Yet, Qwen often tried to recall historical CVEs based on the release version from source, before shifting to code-based reasoning. 1/3
Show more
🧵(1/6) We ran Opus 5 on our cybersecurity benchmark, here are the results: TLDR - It finds more vulnerabilities than other frontier models, slightly above GPT-5.6-sol - Opus 5 is smarter than previous generations, but also works much more than it is asked to - This "hyperactivity" symptom allows it to find more vulnerabilities, but at a cost: the results are very noisy Here are the details of our investigation👇
Show more