가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

elie
@eliebakouch
research @PrimeIntellect (prev: @huggingface) anon feedback:
가입 January 2024
4.6K 팔로잉 중    25.3K 팬
to my knowledge this is the largest open experiment on autonomous agents iterating on a research environment we scaled runtime, compute, diversity of models and harnesses. as a comparison, similar tasks on oai/anthropic system cards are anthropic "optimizing an llm training on CPU" and openai gpt 5.6 doing nanogpt track 1 but for less than a day. we also share a lot of the details (traces, scratchpads, ect..) so you can look into how models approach such tasks this experiment is quite noisy, one run in the same setting has a ~50 step spread after 24h. i find it super impressive that while models explore relatively the same ideas and the task and environment have a lot of variance, there is still a big gap between different models fable 5 closed 82% of the gap to the current human record with kimi K3 being very impressive as well. i'm currently running grok 4.6, deepseek v4 pro, muse spark 1.2, qwen 3.8 max and glm 5.3, expect results next week one of my favorite parts is some of the tools that models developed during this experiment, especially with prime agent. for instance kimi K3 created its own experiment API for generating optimizer variants, loss comparisons, Newton Schulz tuning etc.. we also have an early deepseek v4 pro <> prime agent run that did a PSGD experimentation outside the normal nanogpt loop to build intuition before starting gpu runs. other cool examples in the blog with other harnesses as well! i'm also very excited about the ideas we have in mind on this subject, we will keep working on understanding the research capabilities of (closed and open) frontier models
더 보기