Register and share your invite link to earn from video plays and referrals.

gongy
@_gongy
inference lead @modal
31 Following    2K Followers
found out we're serving 40% of Kimi K3 traffic on OpenRouter ... the guy working on it is tired. anyone wanna help us work on inference?
I sat down with @silasalberti, Head of Research @cognition, to chat about the intersection of RL and inference -- from training frontier coding models to running inference at scale. Watch till the end for bonus doggo! :o 0:00 — Intros 1:04 — What's hardest to get right in an RL run 3:39 — How Cognition started using Modal 4:03 — What to weigh when buying inference 4:58 — The latency/throughput Pareto frontier 7:14 — DFlash reveal (!!) 8:44 — Why faster runtime matters 9:56 — Memory-bound vs. compute-bound 10:57 — Tree-based speculation 14:20 — Hot take: RL and inference optimization converging 15:28 — Online training of the speculator 16:22 — "Auto inference" 17:27 — Agents competing on optimization problems 18:27 — What inference providers still get wrong
Show more
Modal has 460 TPS for Kimi K3 on release day!