A model is only as good as its data, and we’ve long since exhausted the internet. From here on out, model progress is gated by data production.
@mercor_ai’s
@BrendanFoody joined us at our Sovereign AI event to talk about how RL environments get built, and why your data might be your real moat:
00:00 Introduction
00:47 A short history of the data market: crowdsourcing to agentic data
02:29 What an RL environment is: worlds, apps, tasks
03:57 Why only humans can measure the frontier
05:35 Building verifiers is the hard part
06:44 Walkthrough: a real legal RL environment
08:18 Leaderboards — and what open weights change
09:45 Post-training results on Apex Agents
11:17 Three ways companies buy data
12:49 Q&A: How do you price data?
14:17 Q&A: What "data quality" actually means
16:42 Q&A: The misunderstanding about synthetic data
18:17 Q&A: Why RL environments now — and what comes after
21:20 Q&A: Can you scale rubric generation with models?
23:00 Q&A: RL environments for cyber defense
25:33 Q&A: Build data in-house or partner?