登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Feitong Yang
@feitong_yang
Building products for scientists @OpenAI · AI for science · opinions my own
参加 September 2018
234 フォロー中    1.2K ファン
Oh wow, it only needed 6 months for model to go from <1% to >99% for solving Arc3. Maybe with RSI, it would only need 3? Making good eval is harder and harder, but making good eval is more and more important.
もっと見る
Side note: when we released ARC-AGI-3 in March, and frontier models scored <1% on it, a few Singularitarian poasters took it as a personal insult, and got very worked up about it. They argued the benchmark was fundamentally broken, that it could not even be solved by the smartest humans, that the max reachable score was actually 40%, etc. We had to deal with a torrent of insults and hate poasts since because we had released an unsaturated benchmark. As it turns out, the benchmark is perfectly calibrated. It is straightforward for a human to score 100% if they do better than average people – all you need is to use fewer actions than our human baseline (which is not a strong baseline, as we used unfiltered human testers). And naturally as a result it's also very feasible for AI to score 100% once real progress towards agentic general intelligence has been made. The trajectory of AI from <1% to 100% over the course of 6 months shows that the benchmark was able to snapshot the recent rise of agentic capabilities. And that rise has happened faster than most people expected, including us.
もっと見る