註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Greg Kamradt
@GregKamradt
President @arcprize, builder/engineer
加入 January 2011
1K 正在關注    51K 粉絲
> one of my favorite parts is some of the tools that models developed during this experiment The next move 37 is is upon us
to my knowledge this is the largest open experiment on autonomous agents iterating on a research environment we scaled runtime, compute, diversity of models and harnesses. as a comparison, similar tasks on oai/anthropic system cards are anthropic "optimizing an llm training on CPU" and openai gpt 5.6 doing nanogpt track 1 but for less than a day. we also share a lot of the details (traces, scratchpads, ect..) so you can look into how models approach such tasks this experiment is quite noisy, one run in the same setting has a ~50 step spread after 24h. i find it super impressive that while models explore relatively the same ideas and the task and environment have a lot of variance, there is still a big gap between different models fable 5 closed 82% of the gap to the current human record with kimi K3 being very impressive as well. i'm currently running grok 4.6, deepseek v4 pro, muse spark 1.2, qwen 3.8 max and glm 5.3, expect results next week one of my favorite parts is some of the tools that models developed during this experiment, especially with prime agent. for instance kimi K3 created its own experiment API for generating optimizer variants, loss comparisons, Newton Schulz tuning etc.. we also have an early deepseek v4 pro <> prime agent run that did a PSGD experimentation outside the normal nanogpt loop to build intuition before starting gpu runs. other cool examples in the blog with other harnesses as well! i'm also very excited about the ideas we have in mind on this subject, we will keep working on understanding the research capabilities of (closed and open) frontier models
顯示更多