yup, spot on! Traces are the biggest asset for teams to improve their agents over time with Evals, Environments, & human digestible reports explaining what users want
with LangSmith we obsess over tooling to help every team understand their agent traces at scale
- online monitoring of every trace (custom models, Jev, super cheap small models, built-in rubrics)
- clustering traces and errors into human digestible views
- turning errors into evals that can be hill-climbed
- and open sourcing tooling to build Harbor Environments for Evals or RL from your data
we’re very “look at the data” pilled, there’s alpha hiding in there and we want to help you find it
more coming soon on helping every team use their valuable Trace data for post-training 👀
Most companies aren't leveraging their most valuable asset - their traces.
There’s a ton of software value to build around raw inference and models like Jev open up a whole new set of possibilities because of their architecture and how cheap they are to run.
When you’re generating billions of tokens across training and production, you need to understand which failures keep happening and how often.
In this example, we use a frontier model on sampled traces to build a failure taxonomy. Then, we freeze it for an annotation pass and use Jev to classify the full corpus. This way, the expensive work of figuring out what to look for doesn’t need to happen on every trace!
For one annotation across 10k traces, our benchmark estimates came out to about $11 with Jev versus $479 with Haiku 4.5. The implication here is that it is a lot more practical to build scalable systems around model observability.
We’re building a bunch of stuff like this in AC2 because we want customers to get more out of their inference.
We are in a world where the number of tokens being produced is increasing exponentially. This only highlights the need for observability infrastructure.