I sat down with
@silasalberti, Head of Research
@cognition, to chat about the intersection of RL and inference -- from training frontier coding models to running inference at scale.
Watch till the end for bonus doggo! :o
0:00 — Intros
1:04 — What's hardest to get right in an RL run
3:39 — How Cognition started using Modal
4:03 — What to weigh when buying inference
4:58 — The latency/throughput Pareto frontier
7:14 — DFlash reveal (!!)
8:44 — Why faster runtime matters
9:56 — Memory-bound vs. compute-bound
10:57 — Tree-based speculation
14:20 — Hot take: RL and inference optimization converging
15:28 — Online training of the speculator
16:22 — "Auto inference"
17:27 — Agents competing on optimization problems
18:27 — What inference providers still get wrong