Register and share your invite link to earn from video plays and referrals.

Julio Pereyra
@ItsJulioPereyra
30 Following    861 Followers
Some harness + training cooptimization we did with the cracked team @baseten. Tons of banger charts but my favorite is data room read coverage which moves from < 1% (base agent) -> 64% (RLM harness) -> 92% (RLM harness + RL) For basic behavioral strategies like "read the whole data room" it still surprises me every day (1) how strong models still don't just know them but, more excitingly, (2) how quickly any model learns them through RL in the right environment!
Show more
APEX-Agents is one of the deepest benchmarks in legal work. @mercor put a ton of effort into realistic worlds with large filesystems and details down to calendars, inboxes and more all generated by human experts actually solving tasks. That gains on synthetic data like LAB generalize to scores on this benchmark was one of our most promising results and its been awesome to partner with Mercor on both benchmarks (and some very cool future ones!)
Show more
Great working with the @EngramLab team to explore model memory as a way to improve personalization and effective knowledge recall We're particularly excited about how models improve trajectories through memory. Rather than just learning to use search tools better, memory enables more directed, efficient, and effective searches that result in far higher intelligence per token. Promising for a future where agents never start from scratch!
Show more
Environments to train agents for diligence is one of my current favorites of our ongoing research projects. A few reasons why: