登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Alex Dimakis
@AlexGDimakis
Professor, UC berkeley | Founder @bespokelabsai |
参加 April 2009
2.8K フォロー中    25.1K ファン
More data from our experiment on agents vs humans on 14 day tasks. We compare human expert coders to coding agents on the same tasks (from AtCoder Heuristic Contest). The exciting finding is that humans scale super-linearly. This is evidence that humans do continual learning, while they are solving a problem! I.e. they learn more about the coding problem they are trying to solve and scale fundamentally better compared to randomly trying things in a memoryless fashion. This result consistently happens for different harnesses and problems. We release our paper on this: One interesting finding as an example: If you're planning to spend 5M tokens for a problem you should have one claude code session. If you have 30M you should have 2 sessions running independently for 15M each, and pick the best. If you have 100M token budget, you should have 3 sessions run independently and pick the best. Obviously communication between agents should do much better than this naive split of independent agents, but this has not been measured in a controlled way, as far as I know.
もっと見る
🧵 With unlimited compute, how fast can agents surpass humans? We introduce Elo-per-token analysis to profile agent performance curves across multiple open-ended tasks. • Agents initially scale faster than repeated sampling, but over long horizons converge toward their theoretical log-linear scaling curve. Humans, in contrast, improve superlinearly. • These curves also tell us how to spend test-time compute: the scaling inflection point gives a simple rule for splitting a fixed budget across agent sessions. Split a long run in a principled way, and you can get significant gains over a single run. • Fitting human-time and agent-token curves also gives us a fun way to translate AI compute into human time. Taking OpenAI’s ~130B-token Navier–Stokes run as input and extrapolating across the two curves gives an equivalent of ~41 years of work by a mathematician at 8 hours/day 😮.
もっと見る