Based on our experiments, Jev bears big potential for evals
✔️ It is up to 400x cheaper than frontier models, making full scoring of production traces feasible, no more sampling needed, even at huge scale
✔️ It allows for concurrent questions on the same state, where LLMs previously struggled with interaction effects and common wisdom was to not mix evals in the same requests. This makes Jev even cheaper.
✔️ The structure of Jev's output forces thinking in distinct categories during setup, improving the engagement with defining the exact evaluator categories
Play around with this demo here:
Jev-powered evals are now available in Langfuse
👉️ Pay 40-400x less than with frontier models
👉️ Run Jev (by @typesafeai) on all production traces without sampling
👉️ Check multiple criteria at minimal extra cost
@marcklingen sat down to demo all langfuse features in a single video.
Agent observability, monitoring, evals, prompt management, and what is special about Langfuse.
Great place to start if you are new, or catch up on latest changes.