TL;DR Swapping an LLM judge for a purpose-built model called Jev in agent evals reportedly delivers 100% accuracy at less than 1/80th the cost.
Title: Jev-as-a-Judge for Agent Evals
URL:
Points
โ๏ธ It's proposed as a fix for a real dilemma: code-based evals are too rigid, LLM-as-a-judge is too non-deterministic
๐ฏ Across 500 repeated judgments, Jev hit 100% agreement, while Claude only reached 80.0%
๐ Jev also had the lowest variance in quality scores; other judges were up to 913x more variable
๐ฐ Evaluating 5 requests cost $0.34 with Jev versus $28.17 with Claude โ a massive gap
โก Average response time was just 0.44 seconds, making it fast as well as cheap
๐ It's framed as a genuinely "third form" of agent evaluator, alongside code-based checks and LLM-as-a-judge
Being able to run judgments cheaply and repeatedly could reshape the whole feedback loop of building agents.
#
AgentEvals# #
LangChain#