Really fun to hang again with my friend 🃏
@polynoamial (OpenAI research scientist, our first guest ever on
@NoPriorsPod in early 2023) to talk about the implications of large test-time compute, and what happens when models are given $10M budgets to spend on a single task. Topics:
01:23 – Why Benchmarks Are Broken
04:19 – Compute Budgets and Projections
06:48 – How Long Should Models Think?
08:01 – Benchmarkmaxxing
09:48 – Noam's Evals
12:40 – Safety (When Model Capability Scales With Spend)
16:09 – Implications For the Model Release Cycle
18:34 – Latent Model Capability
22:27 – Limits on Recursive Self-Improvement
28:38 – Large-Scale Multi-Agent Coordination
30:39 – Competition at the Frontier
33:19 – Breaking the Benchmark Grid Equilibrium
34:57 – Why Benchmarks Should be Scaled by Cost