Topics clusters every production trace to surface what agents are doing. But 100% coverage only works with a small model that clears the quality bar.
So we worked with
@baseten to build a benchmark from real Topics traces, iterated against the failure modes that mattered most, and found the model that met our needs.
Anyone running a model over production data for their own product can do the same.
Read more →