Just updated our AI evals FAQ with 5 new questions, 48 questions & answers total!
New FAQs just added:
- Do I need a reference answer or rubric before annotating data?
- How can I do evals when traces contain sensitive data?
- How do you review a trace that is really large?
- How much context should I give a LLM judge?
- What should I do when my “gold” eval dataset becomes stale?
It's all here: