Register and share your invite link to earn from video plays and referrals.

Seth Karten
@sethkarten
Prime Agent | Continual Harness, LLM Economist | Research @PrimeIntellect | PhD @Princeton | NSF GRFP
608 Following    3.9K Followers
prime agent v0.9.5 we fixed a lot of bugs and, of course, we had prime agent feature its favorite updates it picked our perf work. then it created the video itself.
Prime Agent reached 20k stars on GitHub! Thanks for being such an amazing community 🦋
My YC Paper Club talk on Prime Agent is out. I talked about moving beyond naive prompting toward an agentic OS, and how eval-driven harness design can expose more of a model’s underlying capabilities through persistent computation, memory, and agent-to-agent communication.
Show more
I asked Prime Agent to make a 30-second video about its favorite changes in v0.9.1. This is what it made.
New paper alert! Check out our arxiv version of Prime Agent. We have added additional commentary on the role and design of our harness almost as the agentic OS layer and some additional results (including our livestreamed factorio runs!) In our roadmap goong forward we will be addressing all comments about short horizon SWE behavior and other requested features.
Show more
Model-harness co-design starts with data. Real world Prime Agent traces can help improve the next generation of open models. You can help by turning on trace collection (default is off). Shared traces will be anonymized and used for building better agents and contributing strong open models and harnesses for the community.
Show more
Prime-Agent playing Factorio
We are running a live eval stress-testing Prime Agent on Factorio
Prime Agent on Factorio! The FLE leaderboard is very outdated (we've seen production scores of 1M+ fairly easily) so I'm really hoping it gets revived :p
We are running a live eval stress-testing Prime Agent on Factorio
We are running a live eval stress-testing Prime Agent on Factorio
One of my main visions for prime agent
Prime-agent is indeed a good general harness for long-horizon tasks GLM 5.2 results on FutureSim Q2
PrimeAgent has now been trending #1# on GitHub for 3 days 👀
The future is multi-agent. Cooperative, decentralized, and with communication. It is great to see so many MARL problems relevant again in the LLM agents era The only question is which eval to test these properties...?
Show more
Today, we’re extending our RL stack beyond individual agents to multi-agent systems. You can now express arbitrary agent interactions and train them.
BREAKING: Prime Agent is trending #1# on github
i set /rlm-max-depth to 10 in prime agent.. the agents yearn to be recursive
Prime Intellect researcher @sethkarten reveals the reason Prime Agent hit 95.5% on ARC-AGI-3 with Opus 5, matching the human expert baseline: "The abstract and symbolic reasoning that is provided by learning how to model the world in these novel game environments in ARC-AGI-3 make it a very useful benchmark for general reasoning capabilities of the models." "95.5% is a fantastic score. That was with Prime Agent running with Opus 5, and we found that Opus 5 was able to greatly take advantage of the capabilities of Prime Agent, more so than the other ones we benchmarked against." "To put in perspective, the human baseline that ARC-AGI released regarding what they expect humans to be able to do in ARC-AGI-3 was 95.4. So we're on par now with potentially saturating the benchmark." @a1zhang @PrimeIntellect
Show more
Really excited to announce Prime Agent, an RLM-native coding TUI! I've been really focused on RLM research and properties of similar harnesses, but ultimately it's all about bringing out stronger capabilities to users. This agent is really good despite not having any models trained around it. I've been using it for my own research with all the shiny new open-weight models and it works great! @PrimeIntellect cooked as always Go try it out :)
Show more
Introducing Prime Agent: A self-improving RLM harness for coding and long-running autonomous tasks. Designed to be both token-efficient and expressive through programmatic tool calling, context as a variable, multi-agent messaging, and a self-modifiable harness state.
Show more
0
37
1.1K
72
Forward to community