A PE operating partner asked us to build production AI agents inside a portfolio company's billing system, processing real healthcare claims under HIPAA.
Two people hand-wrote every rule in their claims engine across 300+ denial codes and payer logic that changes quarterly.
Four months later, seven production agents handle it with zero patient data exposure.
First month, we didn't touch a model. We mapped their data: where it sits and what's missing, so agents reason from structured facts instead of guessing.
I've watched teams skip this step across dozens of engagements. They bolt a model onto the product, watch it hallucinate over unstructured inputs, and decide AI isn't ready for their industry. The data work is what makes it ready.
We built an enrichment layer that assembles 34 dynamic variables per claim before any LLM sees it, pre-computed and versioned so the agent receives ranked facts instead of searching for context.
Every agent follows one pattern: pre-compute context, strip all patient data before the model sees it, validate output against a strict schema, let deterministic code accept or reject the action. If the output falls outside the allowlist, the system fails closed.
Seven agents, each locked to a single workflow like denied claim follow-up or billing reconciliation, each running its own enrichment payload.
Then we built the eval harness.
Every agent runs against a curated test suite before any update reaches production. When a model provider ships a new version or payer logic changes, the harness catches regression before a single live claim is affected. The flagship agent reconciles denials to the penny: 59 out of 60 on the eval set.
Most teams launch an agent and hope it keeps working. We launch one and prove it does on every deployment.
We route calls across two model providers. Swapping one changes nothing in the output because the eval harness verifies it.
Model integration was the shortest line item in the four-month build.
The operating partner now benchmarks the rest of the portfolio against this system.
That's the line between a portfolio company running AI and one still running demos.