Later this month I'll be hosting two mini-workshops on the skills that I think will differentiate the best product teams from the rest.
AI Evals: The New Discovery Habit
🗓️ September 23, 2026, 9am-10:30am PDT
In this session, you'll get introduced to what AI evals are, you'll receive a blueprint for how to get started with this new skill, and we'll leave ample time for your questions. We'll explore both how you can use evals in your personal workflows and in your customer-facing products and features.
Story-Based Customer Interviews
🗓️ September 24, 2026, 9am-10:30am PDT
As delivery gets cheaper and faster, it becomes more tempting to simply build every idea. Choosing what to build is becoming more critical than ever. In this session, you'll get introduced to story-based customer interviewing. You'll learn why continuously collecting customer stories is the secret to uncovering opportunities that can guide what we build next. And you'll get a round of hands-on practice with this new skill.
Both workshops are free for Supporting Members and CDH Members of Invites go out to all subscribers one week before the event. Get access by subscribing today.
To subscribe, go to On desktop, click on the Subscribe button in the top right. On mobile, the Subscribe button is at the bottom of the hamburger menu.
Show more
🎙️Facilitation Skills
Conflict on a product team isn't a warning sign — the absence of it usually is.
In this episode of All Things Product, Petra Wille and Teresa Torres tackle a question Teresa hears constantly: what should a product trio do when collaboration breaks down and disagreement turns into friction?
Petra makes the case for something many teams quietly abandoned — the retrospective — and argues that facilitation is a genuine, learnable skill that product people should invest in, not outsource by default. Teresa introduces a framework from academic research on team dynamics: the distinction between task conflict (we disagree on how to do the work) and relationship conflict (something about how we work together puts us at odds). The two require completely different responses, and confusing them is where teams get stuck.
Together they get practical about team charters, when to bring in a neutral third party, how to choose that person, joint escalations, and why running an experiment often beats arguing over opinions.
If your trio is avoiding hard conversations — or having the same one on repeat — this one's for you.
Key takeaways
🔄 Retrospectives still work. Agile may feel unfashionable, but the practice of pausing to ask "how are we working together?" is one of the most valuable things a team can protect time for.
🎤 Facilitation is a skill, not a personality trait. Someone in the organization needs to be good at it — and product people should be learning it, not just relying on an agile coach or scrum master.
📜 Team charters go deeper than mission statements. The powerful version isn't "what's our outcome" — it's working hours, communication preferences, and how each person likes to make decisions. Making implicit norms explicit prevents friction before it starts.
⚖️ Task conflict vs. relationship conflict. Task conflict is disagreement about the work. Relationship conflict is about how you work together. Team charters address the second; shared discovery addresses the first.
👍 Task conflict is good. It's how teams get to better problem solving. If there's none at all, that's likely a psychological safety problem, not a harmony win.
🗣️ Bring task conflict to the team. What looks like two people disagreeing is often a disagreement others share silently — or one someone else on the team can resolve.
🤝 Relationship conflict starts one-on-one. Try to resolve it directly first; HR or mediation is the escalation path, not the opening move.
🎯 Choose your facilitator deliberately. Ask whether context helps or hurts. An embedded agile coach may know too much history to stay neutral; sometimes the product person — or a senior engineer with strong facilitation instincts — is the better pick.
🚪 Leave your ego at the door. A facilitator writes the headline post-its and nothing else. Skilled facilitators even signal role changes physically — Petra describes one who literally switched chairs to mark "team member" vs. "facilitator."
🧪 When it's opinion vs. opinion, run an experiment. Design a test rather than escalating a debate.
📝 Joint escalation as a tool. Having both parties write down the conflict together often shrinks it — and gives leadership something concrete when escalation is needed.
Links to Spotify, Apple Podcasts, and YouTube:
📺Youtube:
🎵Spotify:
🍎Apple Podcast:
Give it a listen and share your thoughts in the comments below.💬👇
Show more
These results are pretty surprising to me.
AI evals have been the "it" skill for product teams for over a year. I've even called evals a new discovery habit.
But I still meet product teams who only have a vague idea of what evals are. And it's not their fault. Most of the writing on this topic is intended for engineers or just isn't specific enough.
I recently created an in-depth eval guide to explain what evals are and why product teams can and should create them. I did my best to make it practical, hands-on, and easy to follow.
AI evals (short for evaluations) are methods for measuring whether an AI product or workflow is performing well. Evals give teams confidence that their AI applications are doing what they expect them to do. They help teams maintain quality and catch issues before they reach users.
Similar to other discovery habits like interviewing and assumption testing, evals can act as a feedback loop to ensure we are on the right track.
If you want to learn more about this new discovery habit, explore my new guide:
Show more
"The only way to know if our AI products and workflows are any good is with evals." 💡
If you're using AI to write PRDs, analyze customer feedback, or build customer-facing AI products, you need to understand evals. This guide breaks down what evals actually are and why product teams should be building them.
Here's what you'll learn:
🔍 What evals are and why they're different from traditional software testing
📊 How to do error analysis to identify which mistakes matter most
🛠️ Four types of evals: golden datasets, code assertions, LLM-as-a-Judge, and customer feedback
✅ How to choose the right eval for each type of error
🔄 How to run experiments and measure if your changes actually work
The key insight: with LLMs, you can't just test once and expect the same result. You need to measure how often your AI gets it right, and that starts with defining what "right" actually looks like for your specific product.
Check out the article:
❓ What's one error you've noticed in an AI tool you use regularly? Share your thoughts in the comments below.
Show more
How does Capital One still not have 2FA?
"I'm going to tell you my story because I think it's an amazing story of continuous improvement, of how teeny-tiny steps compound over time."
Over the past year, my work has transformed completely—and it all started with a broken ankle and a willingness to get curious about AI. In this article, I share how I went from a skeptic to building multiple AI products, even though I had a limited engineering background.
Here's what you'll learn:
🔧 How small experiments can snowball into real products (I launched my first AI tool just three weeks after starting to experiment)
🧠 Why AI evals became the missing discovery habit—a feedback loop that helped me continuously improve my products
🎯 How personal productivity experiments taught me the skills I needed to build production AI products
🤝 Why I still believes in talking to customers, but see new opportunities for AI to be additive rather than replaceable
💡 The hero's journey framework I used to make sense of my transformation
🚀 Real examples of the products I built: Interview Coach, Business Fundamentals Coach, AI-generated interview snapshots, and more
The core message? You don't have to be an engineer to figure this out. You have access to expert tutors 24/7, and teeny-tiny steps compound over time.
Read the article (or watch the talk):
❓ What's one small experiment or skill you've been curious about trying but haven't started yet? Share your thoughts in the comments below.
Show more
True
There is a right way and a wrong way to learn from the people you want to serve.
"It can feel like you are the lone champion pushing for change in your organization."
That's why we are reading Continuous Discovery Habits together throughout 2026. 📚
The book turns five this year and many of you have bought and read it. But reading isn't the same as doing. So we are reading together—one section per month—with discussion questions, practical exercises, and resources to help you actually build the habits.
Here's what you'll get each month:
📖 Monthly reading guides with reflection questions and exercises
🎥 Short videos you can share with teammates to spread the ideas
💬 Quarterly live discussion sessions to connect with other practitioners
By December, you won't just understand continuous discovery—you'll be practicing it.
This month's reading covers Chapter 9—it's all about deconstructing your product ideas into their underlying assumptions. You'll learn about story mapping, pre-mortems, assumption testing, and more
Read the chapter:
🤔 Does your team break your ideas down into their underlying assumptions? Share your experience in the comments.
Show more
Couldn’t get starlink to work after talking to two real humans. I guess we’ll just take the day off. First night in our new rig.
Trying to setup a starlink mini. When I get to the choose a plan step, it says no plans available. Starlink is available in my area and I’m trying to buy a roam subscription. Anyone else run into this?
Their AI phone help said I have to wait until it’s available in my area. It is available in my area.
Show more
Much of this analysis can also apply to Leo Carlsson. The Philly offer sheet certainly came with sticker shock, but depending on how you project Carlsson's growth trajectory, he very well may grow into his contract. Carlsson at 21 didn't have as strong of a season as Celebrini at 19. But he certainly belongs in the top 10 1C category already.
If I were a Ducks fan, I'd be more comfortable if his compensation fell more in the middle of the pack of the top 10 1Cs. But we can't always get what we want when offer sheets are invovled. Also, Carlsson's contract starts 1 year sooner than Celebrini, so that pushes things a little bit.
My take: Carlsson wont' live up to 100% of his contract value, but the mismatch isn't as extreme as much of the media presents.
Show more
Doctors juggle a packed schedule, ten minutes per patient, and constant context-switching—while being expected to be the expert on absolutely everything.
Hertility Health's AI assistant helps close that gap: pinpointing the most relevant information about each patient and suggesting a diagnosis and follow-up care doctors might otherwise miss. It's a second opinion for the doctor, and a way for the patient to finally feel taken seriously.
👉 Find a link to the full episode here:
Spotify:
Apple Podcast:
YouTube:
Show more
NHS clinicians get ten minutes with a generalist, if they can get an appointment at all—leaving little room for the specific, personal questions women's health actually requires.
Hertility Health made a longer intake mandatory before purchase, and something unexpected happened: conversion went up. Women finally felt heard before being handed a test kit—and clinicians got better information to work with, easing the pressure that drives burnout.
👉 Find a link to the full episode here:
Spotify:
Apple Podcast:
YouTube:
Show more
Endometriosis can take up to ten years to diagnose—and even then, women often have to keep fighting for the ongoing care they need.
Hertility Health built GynAI to close that gap. Every data point—an online health assessment, blood test results, scan images, even a clinician conversation—stacks into a clearer likelihood score, so doctors can refer women onward with confidence, much earlier in the process.
👉 Find a link to the full episode here:
Spotify:
Apple Podcast:
YouTube:
Show more
How do you build trustworthy AI diagnostic tools in one of medicine's most historically under-researched areas?
In this episode of Just Now Possible, Teresa Torres talks with Tulsi Patel (Director of Product and Technology), Lorna Brightmore (Head of Data and AI), and Jack Pickard (Head of Engineering) at Hertility, a UK and Ireland-based women's health tech company. Hertility combines an in-depth online health assessment with at-home hormone testing and clinician-reviewed reports to help diagnose conditions spanning menstruation to menopause.
Built on seven years of data linking symptoms, blood results, and pelvic ultrasound scans for over a million women, the team walks through two AI products in development: a Bayesian network that gives clinicians probability-based diagnoses instead of binary calls, and a scan automation pipeline that classifies ultrasound images, measures follicle counts and ovarian volume, and drafts clinical letters using an agentic loop that checks its own output against patient data before a human ever reviews it.
You'll hear how the team guards against automation bias, builds clinician trust through transparency, minimizes PII before it ever reaches a model, and treats healthcare regulation as a design constraint from day one rather than a last-minute scramble. It's a detailed look at what it takes to bring AI into one of the most sensitive, tightly regulated corners of healthcare.
Guests:
- Tulsi Patel – Director of Product and Technology, Hertility
- Lorna Brightmore – Head of Data and AI, Hertility
- Jack Pickard – Head of Engineering, Hertility
What we cover:
- What makes Hertility's data set unique: seven years of linked symptoms, blood tests, and pelvic scans from over a million women
- How uses a Bayesian network to give clinicians probability-based diagnoses instead of binary yes/no calls
- Why showing clinicians the reasoning behind a diagnosis—not just the label—builds trust and speeds up triage
- Guarding against automation bias with holdout sets and independent, fresh-eyes review
- Inside the scan automation pipeline: classifying ultrasound images, detecting follicles, and measuring ovarian volume more precisely than manual methods
- Using an agentic loop to check AI-drafted clinical letters against patient data and catch hallucinations before a human sees them
- The infrastructure challenge of securely piping DICOM ultrasound images from third-party scan providers into Hertility's systems
- How Hertility handles PII and PHI: pseudonymization, data minimization, and running models in-house on AWS Bedrock
- Why treating healthcare regulation as a product requirement from day one makes AI products more scalable, not slower
Key Takeaways:
- Probabilistic, transparent AI outputs build more clinician trust than binary classifications.
- Guardrails against automation bias are as important as the model itself.
- Data minimization and in-house infrastructure make it possible to build AI responsibly with sensitive health data.
- Treating regulation as a design constraint from day one makes AI products more defensible and scalable, not slower.
Resources & Links:
- Hertility — At-home hormone testing and reproductive health diagnostics for women in the UK and Ireland
- AWS Bedrock — The platform Hertility uses to run LLMs in-house under its own governance and regulatory controls
- PyTorch — The foundation for Hertility's in-house image classification and contouring models
Chapters:
00:00 Meet the Team
00:13 What Hertility Does
01:51 How Customers Access It
04:06 A Unique Women's Health Dataset
07:03 Mission and Efficiency with AI
10:03 Why Long Assessments Convert
13:52 Before AI Workflows
16:52 Research Publications and Impact
18:48 GynAI Reducing Time to Diagnosis
21:21 Triage and Clinician Support
24:37 Keeping Patient UX the Same
26:12 Bayesian Network and Explainability
30:19 Multiple Diagnoses and Probabilities
32:37 Probabilistic Diagnosis Shift
33:50 Clinician Adoption and Workflow Fit
34:58 Communicating Medical Uncertainty
36:43 Scan Automation Overview
40:30 In House Image Analysis
44:25 DICOM Pipeline Engineering
47:30 Evals and Automation Bias
50:31 LLM Letter Guardrails
56:47 PHI Handling and Regulations
01:00:43 Infrastructure Choices and Wrap Up
Listen on Spotify, Apple Podcasts, or watch on YouTube.
Spotify:
Apple Podcast:
YouTube:
Show more
I'm seeing a huge difference in performance (in a bad direction) across my evals moving from Sonnet 4.6 to Sonnet 5. Previously, my service ran at temperature 0. I'm wondering what prompt strategies people are using to constrain Sonnet 5 now that temperature is no longer an option.
Show more
All Things Product with Teresa Torres and Petra Wille is on summer break. New episodes will return on Tuesday, September 8th.
If you are looking for something else to listen to, we have a back catalog of 68 episodes. And if you are looking to mix it up, try my other podcast Just Now Possible where I interview product teams about the AI products they are building.
All Things Product:
Just Now Possible:
Show more