Register and share your invite link to earn from video plays and referrals.

Hamel Husain
@HamelHusain
Evals Evals Evals - About Me:
2.7K Following    56.9K Followers
This is the most valuable free resource we've created on AI Evals 🎉 (not exaggerating!) I organized our public materials into this guide. It allows you to find answers to your eval problems w/o searching aimlessly. Humans: pick the row that sounds like you. Agents: point it at the post The material draw on 60+ hours of office hours from our Evals course, where @sh_reya and I have taught 5k engineers and PMs Evals. We add new material often, with 15 FAQs added in the last two weeks. Recent additions are marked with a "New" badge. You can find all of these and more here:
Show more
0
50
914
106
Forward to community
“I used my agent to formally verify my software and it fixed all these bugs” is the new “I asked an LLM judge to fix my LLM output”. The devil is in the details and it is insanely difficult to get the formalisms “right” (eg interpretable, expressive, extensible) for humans and agents to co-create software. Seems that we are in for an entirely new age of learning how to build software
Show more
Very exciting to see how much this post resonated with the community. The uncomfortable dialectic is that (1) evals are necessary to build good agents, but also (2) agents are extremely helpful (and necessary) for automating tedious parts of the evals lifecycle. Human attention and labeling simply cannot scale up to the high amount of unstructured info that is agents. Unfortunately nobody has cracked the most perfect and efficient workflow to do evals. We share our thoughts on our current iteration in this post!!
Show more
I recommend grabbing these skills that @sh_reya and I created and at the very least doing an /eval-audit of your existing pipeline. We've found that people often find low hanging fruits!
Show more
Pro tip: Install this new evals skill from @HamelHusain and @sh_reya, it'll save you many hours and a lot of mistakes
Is Slop Dead? (Testing Writing With 5.5 Opus)
Good Read - "Bulding eval systems that improve your AI product" by @HamelHusain and @sh_reya
Well worth the read! Eval skills are quickly becoming a requirement for most AI jobs.
Pro tip: Install this new evals skill from @HamelHusain and @sh_reya, it'll save you many hours and a lot of mistakes
0
19
1.1K
70
Forward to community
Q: What should you do when your "gold" eval dataset becomes stale? A: Use regular error analysis to find new problems. Keep updating your evals as your product and users change.
Show more
Also fun fact about Lenny's newsletter now that I've written two guest posts there. The quality bar is **insanely high** Lenny rigorously vets each post. These go through many rounds of writing and technical review with world class staff. You have to sweat it out as a writer but its worth it. Posts and ideas get rejected if they don't meet the bar. It's no wonder that Lenny is at the top of the game. Many people over-index on "quantity" but he's shown dedication to quality. There is something to learn if you are a creator/educator in the space. While we are on the topic of Lenny, he is not an overnight success. He has been doing this with a relentless commitment for over a decade 🤯 and the focus pays off. Truly long-term thinking and playing the infinite game on his part.
Show more
Also fun fact about Lenny's newsletter now that I've written two guest posts there. The quality bar is **insanely high** Lenny rigorously vets each post. These go through many rounds of writing and technical review with world class staff. You have to sweat it out as a writer but its worth it. Posts and ideas get rejected if they don't meet the bar. It's no wonder that Lenny is at the top of the game. Many people over-index on "quantity" but he's shown dedication to quality. There is something to learn if you are a creator/educator in the space. While we are on the topic of Lenny, he is not an overnight success. He has been doing this with a relentless commitment for over a decade 🤯 and the focus pays off. Truly long-term thinking and playing the infinite game on his part.
Show more
This post was a labor of love! We’ve distilled thousands of hours of work on AI Evals into a 30 min read + skills you can use to quickly uncover errors in your product. BTW in addition to high quality content, Lenny effectively **pays you** in AI credits to subscribe 🤯 It's really good!
Show more
Evals have been coming up more and more in my conversations with podcast guests and PM friends. Nearly half of the 25 awesome PM job openings I shared last week ask for experience writing evals. And leading companies keep sharing what investing in evals bought them: — @tryramp took its automatic receipt collection from 35% to 83% accuracy. — @Shopify shipped an AI workflow builder that's 2.2x faster and 68% cheaper than the frontier-model setup it replaced. — @harvey__ai rebuilt its AI contract reviewer, nearly doubling its internal quality score. — @cursor_ai tuned its Auto Balance routing, with much higher user satisfaction at 41% lower cost. So I asked the 🐐s of evals, @HamelHusain and @sh_reya, to write an advanced sequel to their very popular "Building eval systems that improve your AI product." Drawing on their work with 50+ AI companies, they share the key step most teams skip, what you should (and shouldn't) automate, and a free plugin that lets a coding agent do most of the heavy lifting. Read it here:
Show more
Jev and the System One Model: RLCD, intelligence/$, reliable AI, & the end of chat-first AI @typesafeai CEO @CompleteSkeptic explains why AI can solve extraordinarily hard problems yet still fail to automate basic work, why Jev is built for reliable decisions inside software instead of chat, why TypeSafe rejects public benchmarks and refusals at the API layer, why data and the right task matter more than brute-force compute, how System One Models could reshape coding agents and software, and why even with $1 billion he wouldn’t pre-train a model from scratch.
Show more
"In practice, this all means that evals and structured experiments are even more important than they are for regular LLM projects." 💯
In the process of migrating the repo I built at Braintrust for @HamelHusain and @sh_reya's AI evals course ( to an all-@pydantic stack: PydanticAI + Logfire + pydantic-evals. Been living in these tools for a while now and the progress on observability and evals is impressive. I'll be blogging and posting findings here as I go, including where I think models like @typesafeai's Jev fit into the end-to-end evals pipeline. Stay tuned ...
Show more
AI-powered operators in SQL are having their day in the sun
Today we're launching prompt_jev(), which brings Jev, a new kind of AI model from @typesafeai, to MotherDuck SQL. Jev makes text classification super fast, dirt cheap, and as easy as prompting an LLM. In our tests, it ran at 50x the speed and 1% the cost of comparable frontier models, with similar or better accuracy. For analytics, this opens up scoring and labeling workflows that used to be too slow or costly. Classify all your data, rerun when your labels change, or filter on meaning in a WHERE clause while doing exploratory data analysis. Live now in MotherDuck!
Show more
People love rage baiting me with comments that it’s not a classifier because it “also produces a confidence score” or it can choose between two options. I’m gonna assume there isn’t brain rot and that it’s just trolling 😬😅 But well played trolls, well played
Show more
What’s wrong with calling it a classifier? Its an established term that describes exactly what this is, along with corresponding literature on how to verify, tune, measure it. Jev is very cool and super useful, because of speed, cost, accuracy and ergonomics. But new terms might hurt more than help. What are your thoughts?
Show more
Who is gonna create general purpose pre trained regression model and call it Rev We already know there is PMF!