๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
๊ฐ€์ž… May 2026
280 ํŒ”๋กœ์ž‰ ์ค‘    420 ํŒฌ
TL;DR Swapping an LLM judge for a purpose-built model called Jev in agent evals reportedly delivers 100% accuracy at less than 1/80th the cost. Title: Jev-as-a-Judge for Agent Evals URL: Points โš–๏ธ It's proposed as a fix for a real dilemma: code-based evals are too rigid, LLM-as-a-judge is too non-deterministic ๐ŸŽฏ Across 500 repeated judgments, Jev hit 100% agreement, while Claude only reached 80.0% ๐Ÿ“‰ Jev also had the lowest variance in quality scores; other judges were up to 913x more variable ๐Ÿ’ฐ Evaluating 5 requests cost $0.34 with Jev versus $28.17 with Claude โ€” a massive gap โšก Average response time was just 0.44 seconds, making it fast as well as cheap ๐Ÿ” It's framed as a genuinely "third form" of agent evaluator, alongside code-based checks and LLM-as-a-judge Being able to run judgments cheaply and repeatedly could reshape the whole feedback loop of building agents. #AgentEvals# #LangChain#
๋” ๋ณด๊ธฐ