Based on our experiments, Jev bears big potential for evals
✔️ It is up to 400x cheaper than frontier models, making full scoring of production traces feasible, no more sampling needed, even at huge scale
✔️ It allows for concurrent questions on the same state, where LLMs previously struggled with interaction effects and common wisdom was to not mix evals in the same requests. This makes Jev even cheaper.
✔️ The structure of Jev's output forces thinking in distinct categories during setup, improving the engagement with defining the exact evaluator categories
Play around with this demo here: