A security benchmark should show the learning loop, not just the winning score.
Krait evolved through 8 methodology versions by running blind against historical smart-contract contests, comparing its output with known results, and turning every miss and false positive into a new pattern, heuristic, or verification gate.
The measured v8 baseline now covers 50 contests at 100% precision with zero false positives per contest. That is a precision result, not a claim that every bug was found—and v8.1/v8.2 are explicitly marked as not yet re-measured.
Inspect the methodology and limitations: