We ran Muse Spark 1.3 on our Cybersecurity benchmark, and it's actually not that good (yet) 😬
- At pass
@1, it rediscovers an average of 19/32 CVEs. In comparison, Grok 4.6 scores 23.3/32
- When pooling the results of 3 runs (pass
@3), Muse scores 24/32. DeepSeek V4 Pro gets 28/32
- The pricing is good compared to many frontier models, but compared to some of the Chinese ones, its use is hard to justify
- We had a cached input token rate of 45%, much lower than we usually get. The model is new, so we decided to include the results as if we had had 90%
Overall, I would say that
@AIatMeta is slowly catching up. The model seems good at coding, but it's definitely not there yet for cybersecurity.
🧵 1/3