Q: What should you do when your "gold" eval dataset becomes stale?
A: Use regular error analysis to find new problems. Keep updating your evals as your product and users change.
Also fun fact about Lenny's newsletter now that I've written two guest posts there. The quality bar is **insanely high**
Lenny rigorously vets each post. These go through many rounds of writing and technical review with world class staff. You have to sweat it out as a writer but its worth it. Posts and ideas get rejected if they don't meet the bar.
It's no wonder that Lenny is at the top of the game. Many people over-index on "quantity" but he's shown dedication to quality. There is something to learn if you are a creator/educator in the space.
While we are on the topic of Lenny, he is not an overnight success. He has been doing this with a relentless commitment for over a decade 🤯 and the focus pays off. Truly long-term thinking and playing the infinite game on his part.
Can you use Jev for Evals? Yes! Remember that a LLM Judge is also classifier*.
Make sure to test your classifiers against human labels and don't overfit.
Hope this helps!
*
Q: What is a trace?
A: The complete record of a user session, from the first query to the final response. It includes the tool calls and intermediate steps.
If you are a Data Scientist, you have never had more alpha than right now.
Data Scientists email me all the time. It is no mistake that I'm focused on evals as a former DS.
I talk more about this here:
Just updated our AI evals FAQ with 5 new questions, 48 questions & answers total!
New FAQs just added:
- Do I need a reference answer or rubric before annotating data?
- How can I do evals when traces contain sensitive data?
- How do you review a trace that is really large?
- How much context should I give a LLM judge?
- What should I do when my “gold” eval dataset becomes stale?
It's all here:
🌶️Codex can do a anything Muse, Bot etc can do quite easily
I get the appeal in the form factor “less is more” but if you are already a codex desktop and are a power user (computer use, remote access, thread management, voice control etc), YAGNI
I still like playing with the other things for educational purposes but also need to avoid tool sprawl unless there is a real benefit