Register and share your invite link to earn from video plays and referrals.

Yash Patil
@ypatil125
Co-Founder, CEO @appliedcompute 🚂 prev: @OpenAI, @Stanford
Joined March 2020
654 Following    11.2K Followers
Most companies aren't leveraging their most valuable asset - their traces. There’s a ton of software value to build around raw inference and models like Jev open up a whole new set of possibilities because of their architecture and how cheap they are to run. When you’re generating billions of tokens across training and production, you need to understand which failures keep happening and how often. In this example, we use a frontier model on sampled traces to build a failure taxonomy. Then, we freeze it for an annotation pass and use Jev to classify the full corpus. This way, the expensive work of figuring out what to look for doesn’t need to happen on every trace! For one annotation across 10k traces, our benchmark estimates came out to about $11 with Jev versus $479 with Haiku 4.5. The implication here is that it is a lot more practical to build scalable systems around model observability. We’re building a bunch of stuff like this in AC2 because we want customers to get more out of their inference. We are in a world where the number of tokens being produced is increasing exponentially. This only highlights the need for observability infrastructure.
Show more
I implemented a system in @appliedcompute’s platform for automated failure mode clustering with Jev to surface errors at an even larger scale than before. RL training produces billions of tokens in traces. I always manually read many traces to understand model behavior, but finding agent failures (like reward hacking / hallucinations) at scale is easy to miss without automation. Here’s how it works:
Show more