In July, 20% of our inference spend was review tables. The most expensive review table queries cost $20k.
Today we’re excited to share work we did with
@appliedcompute that will reduce this cost by 50% by post training a review table model.
Our review table product allows layers to upload thousands of contacts or emails and extract or analyze flexible fields.
Historically we optimized this system via prompt caching, packing (multiple fields per model call), routing, batching, etc.
Once we hit the limits of gradient free approaches we wanted to see if post training could improve on an already strong baseline.
Review tables turned out to be a perfect candidate for post training because verification is easy and we’ve built large datasets to do so.
Awesome work by
@vtrengarajan, our vault and labs team as well as applied compute. Detailed write up below.