登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

tenex
@tenex_labs
We are your full-stack AI partner helping you set & execute your enterprise AI strategy at startup speed.
参加 April 2025
9 フォロー中    7.9K ファン
One of our best lessons yet. This is THE only guide you need to evals in 2026.
McKinsey surveyed 2,000 companies in 2025. 51% said AI backfired on them. Top reason? Inaccuracy. From what I can tell, most of these systems weren't broken. They were unreliable. And unreliable is wayyyy worse because you can't predict when it fails. So I got @ashtilawat (the Mr. Miyagi of teaching AI) from @gauntletai to walk me through the solution. Here's his 2026 framework for evaluating if your AI is trustworthy, reliable, and production-ready: 1. build your golden set Identify 30–50 core requests your AI must handle correctly. The stuff that, if broken, makes the whole system useless. And sit with the person whose job this AI is doing/automating/replacing/helping with. 2. test the weird stuff Your golden set covers common requests. But in production, users don't only ask common requests. So build a matrix of categories (topic x complexity) and fill the gaps. Every gap is a corner where failures can hide behind. 3. build a replay harness Record the exact state of every interaction so you can test prompt changes without burning API calls. Think of it like game film... you don't put players back on the field just to review the play. 4. create your rubric Use an LLM to grade outputs on accuracy, completeness, and tone. But calibrate it first -> run 50–100 examples through human and LLM scoring, find disagreements, fix the rubric, repeat until they match. 5. run experiments New model? Prompt rewrite? Run your eval suite against both versions. Ship if the golden set passes, no regressions, and the cost is acceptable. The teams still running production AI on vibes will be f***** in 2026. But the teams building eval libraries are compounding an advantage that gets harder to catch every month. Competitors can copy your product. They can't copy your test cases. h/t @Austen for helping put this together. Full playbook + vid below 👇
もっと見る