11 LLM eval methods AI engineers must know:
(bookmark this)
The tricky part about LLM evaluation is that there is no single metric that tells you whether a system is good.
The right evaluation method depends on what you are trying to measure. You may want to compare an output against a known answer, judge whether it is semantically correct, inspect how an agent behaved across multiple steps, or block unsafe outputs before they reach the user.
A useful way to organize them is by what they are actually evaluating.
๐ฅ๐ฒ๐ณ๐ฒ๐ฟ๐ฒ๐ป๐ฐ๐ฒ-๐ฏ๐ฎ๐๐ฒ๐ฑ ๐ฒ๐๐ฎ๐น๐, ๐๐ต๐ฒ๐ป ๐ด๐ฟ๐ผ๐๐ป๐ฑ ๐๐ฟ๐๐๐ต ๐ฒ๐
๐ถ๐๐๐
โ BLEU measures n-gram precision, checking how much of the generated text overlaps with the reference.
โ ROUGE focuses more on recall, measuring how much of the reference content appears in the output.
โ BERTScore compares contextual embeddings instead of exact words, so semantically similar answers can still score well.
๐๐๐ฑ๐ด๐ฒ-๐ฏ๐ฎ๐๐ฒ๐ฑ ๐ฒ๐๐ฎ๐น๐, ๐๐ต๐ฒ๐ป ๐ด๐ฟ๐ผ๐๐ป๐ฑ ๐๐ฟ๐๐๐ต ๐ถ๐ ๐๐ป๐ฎ๐๐ฎ๐ถ๐น๐ฎ๐ฏ๐น๐ฒ ๐ผ๐ฟ ๐ถ๐ป๐ฐ๐ผ๐บ๐ฝ๐น๐ฒ๐๐ฒ
โ G-Eval uses an LLM to score an output against defined criteria.
โ LLM-as-Judge gives a model a rubric and asks it to score or compare outputs.
โ LLM Juries run multiple independent judges and aggregate their verdicts, reducing dependence on a single evaluator.
๐๐๐บ๐ฎ๐ป ๐ฎ๐ป๐ฑ ๐ฑ๐ฒ๐๐ฒ๐ฟ๐บ๐ถ๐ป๐ถ๐๐๐ถ๐ฐ ๐ฒ๐๐ฎ๐น๐
โ Human Evaluation relies on people to score outputs across dimensions such as correctness, relevance, and helpfulness.
โ DAG-based Evaluation encodes evaluation as deterministic decision logic, routing outputs through explicit checks until a final verdict is reached.
๐๐๐ฎ๐น๐ ๐ฏ๐๐ถ๐น๐ ๐ณ๐ผ๐ฟ ๐ฎ๐ด๐ฒ๐ป๐๐
โ Trajectory Accuracy evaluates the sequence of actions an agent took, not just the final answer.
โ Multi-turn Evaluation scores behavior across an entire conversation, including consistency, memory, and coherence across turns.
๐๐๐ฎ๐น๐ ๐๐ต๐ฎ๐ ๐ฟ๐๐ป ๐ฎ๐ ๐ด๐ฎ๐๐ฒ๐
โ Safety Evaluation checks for things such as toxicity, bias, sensitive-data leakage, and policy violations before an output is accepted.
The important part is that these methods are complementary.
A production agent may need reference-based evals for correctness, judge-based evals for subjective quality, trajectory evals for tool use, and safety evals before anything reaches the user.
If you are building this evaluation layer, Cometโs Opik brings tracing, debugging, test suites, and agent evals into one open-source stack.
Check it out on GitHub:
(don't forget to star ๐)
I also wrote a deeper breakdown of how evaluation fits into a production observability and self-repair loop.
The full article is quoted below.
๋ ๋ณด๊ธฐ