๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Akshay ๐Ÿš€
@akshay_pachaar
Simplifying LLMs, AI Agents, RAG, and Machine Learning for you! โ€ข Co-founder @dailydoseofds_โ€ข BITS Pilani โ€ข 3 Patents โ€ข ex-AI Engineer @ LightningAI
๊ฐ€์ž… July 2012
501 ํŒ”๋กœ์ž‰ ์ค‘    290K ํŒฌ
11 LLM eval methods AI engineers must know: (bookmark this) The tricky part about LLM evaluation is that there is no single metric that tells you whether a system is good. The right evaluation method depends on what you are trying to measure. You may want to compare an output against a known answer, judge whether it is semantically correct, inspect how an agent behaved across multiple steps, or block unsafe outputs before they reach the user. A useful way to organize them is by what they are actually evaluating. ๐—ฅ๐—ฒ๐—ณ๐—ฒ๐—ฟ๐—ฒ๐—ป๐—ฐ๐—ฒ-๐—ฏ๐—ฎ๐˜€๐—ฒ๐—ฑ ๐—ฒ๐˜ƒ๐—ฎ๐—น๐˜€, ๐˜„๐—ต๐—ฒ๐—ป ๐—ด๐—ฟ๐—ผ๐˜‚๐—ป๐—ฑ ๐˜๐—ฟ๐˜‚๐˜๐—ต ๐—ฒ๐˜…๐—ถ๐˜€๐˜๐˜€ โ†’ BLEU measures n-gram precision, checking how much of the generated text overlaps with the reference. โ†’ ROUGE focuses more on recall, measuring how much of the reference content appears in the output. โ†’ BERTScore compares contextual embeddings instead of exact words, so semantically similar answers can still score well. ๐—๐˜‚๐—ฑ๐—ด๐—ฒ-๐—ฏ๐—ฎ๐˜€๐—ฒ๐—ฑ ๐—ฒ๐˜ƒ๐—ฎ๐—น๐˜€, ๐˜„๐—ต๐—ฒ๐—ป ๐—ด๐—ฟ๐—ผ๐˜‚๐—ป๐—ฑ ๐˜๐—ฟ๐˜‚๐˜๐—ต ๐—ถ๐˜€ ๐˜‚๐—ป๐—ฎ๐˜ƒ๐—ฎ๐—ถ๐—น๐—ฎ๐—ฏ๐—น๐—ฒ ๐—ผ๐—ฟ ๐—ถ๐—ป๐—ฐ๐—ผ๐—บ๐—ฝ๐—น๐—ฒ๐˜๐—ฒ โ†’ G-Eval uses an LLM to score an output against defined criteria. โ†’ LLM-as-Judge gives a model a rubric and asks it to score or compare outputs. โ†’ LLM Juries run multiple independent judges and aggregate their verdicts, reducing dependence on a single evaluator. ๐—›๐˜‚๐—บ๐—ฎ๐—ป ๐—ฎ๐—ป๐—ฑ ๐—ฑ๐—ฒ๐˜๐—ฒ๐—ฟ๐—บ๐—ถ๐—ป๐—ถ๐˜€๐˜๐—ถ๐—ฐ ๐—ฒ๐˜ƒ๐—ฎ๐—น๐˜€ โ†’ Human Evaluation relies on people to score outputs across dimensions such as correctness, relevance, and helpfulness. โ†’ DAG-based Evaluation encodes evaluation as deterministic decision logic, routing outputs through explicit checks until a final verdict is reached. ๐—˜๐˜ƒ๐—ฎ๐—น๐˜€ ๐—ฏ๐˜‚๐—ถ๐—น๐˜ ๐—ณ๐—ผ๐—ฟ ๐—ฎ๐—ด๐—ฒ๐—ป๐˜๐˜€ โ†’ Trajectory Accuracy evaluates the sequence of actions an agent took, not just the final answer. โ†’ Multi-turn Evaluation scores behavior across an entire conversation, including consistency, memory, and coherence across turns. ๐—˜๐˜ƒ๐—ฎ๐—น๐˜€ ๐˜๐—ต๐—ฎ๐˜ ๐—ฟ๐˜‚๐—ป ๐—ฎ๐˜€ ๐—ด๐—ฎ๐˜๐—ฒ๐˜€ โ†’ Safety Evaluation checks for things such as toxicity, bias, sensitive-data leakage, and policy violations before an output is accepted. The important part is that these methods are complementary. A production agent may need reference-based evals for correctness, judge-based evals for subjective quality, trajectory evals for tool use, and safety evals before anything reaches the user. If you are building this evaluation layer, Cometโ€™s Opik brings tracing, debugging, test suites, and agent evals into one open-source stack. Check it out on GitHub: (don't forget to star ๐ŸŒŸ) I also wrote a deeper breakdown of how evaluation fits into a production observability and self-repair loop. The full article is quoted below.
๋” ๋ณด๊ธฐ