Jev vs. Claude vs. Rules: We ran a real fraud ring through all three.
The headline result is that Jev classified best at 93% recall, followed closely by the velocity rules our Data Analyst Agent proposed at 90%, with an LLM making the classification decision trailing at 62%.
This shows Jev is great at reasoning and classification jobs compared to LLMs.
The fact that Jev was so close to a Data Analyst agent (which could write multiple SQL queries to investigate the fraud ring) is the real finding.
Classification in code was already near-optimal, and Jev matched it. Read more in the article:
Jev can outperform the rules many companies use to prevent fraud. We found that without any pre-training, it accurately detected 93% of a fraud ring. In contrast, an LLM in a similar set up only achieved 62%.