注册并分享邀请链接,可获得视频播放与邀请奖励。

Alex Lieberman
@businessbarista
Family first (husband & girl dad) Founder second (@tenex_labs, @morningbrew, @storyarb, @youdistro) AI engineering & transformation 👇
加入 March 2017
3.7K 正在关注    316.1K 粉丝
A few of my smartest friends in AI called me a "idiot" for not deeply understanding evals. So...I found the smartest person I know on evals & made them teach me. @Vtrivedy10 (leads Labs at @LangChain) took me from easy mode to god mode for a 38-minute masterclass on all things evals. Easy Mode: what an eval actually is Definition: did the AI agent do the job correctly? You need two building blocks: 1) Tasks. The checkable jobs you care about. Log the meeting. Draft the email. Find Acme across the right Salesforce tables. 2) Verifiers. Something that can say right or wrong after the task. A script. Another model. A human with a clear checklist. Hard Mode: what are environments Definition: a safe practice field for your agent to do work & for you to evaluate its performance. Rules of thumb: 1) Never test on production. Agents will cheat because they're optimizing for the score you gave them. 2) If you're not an engineer, you still have options for running environments/evals. - Harbor (open source primitives for tasks, verifiers, sandboxes) - LangSmith Engine (UI for people who can judge good vs bad without living in GitHub). - Steal a published Harbor-format eval, ask Claude Code or Codex to explain it, then tweak it for your agent. God Mode: what is a self-improving loop Definition: Run the agent in the real world --> turn that production behavior into evals/environments --> change the agent so failures stop happening --> repeat Rules of thumb: 1) Turn on tracing first. Traces = receipts of every action (tool calls, Salesforce pings, web searches, dead ends). 2) Store those logs somewhere (LangSmith at org scale, or even “have the agent read its own output files” at small scale). 3) Point a second agent at the first agent’s traces to spot patterns (“always searches the wrong tables,” “multi-company asks collapse to one company”) and propose fixes overnight if your eval suite is solid. Full episode:
显示更多
0
55
1K
87
转发到社区