I implemented a system in
@appliedcompute’s platform for automated failure mode clustering with Jev to surface errors at an even larger scale than before.
RL training produces billions of tokens in traces. I always manually read many traces to understand model behavior, but finding agent failures (like reward hacking / hallucinations) at scale is easy to miss without automation. Here’s how it works: