Introducing Prime Sandboxes:
MicroVM sandboxes purpose-built for RL training.
Model training requires running tens of thousands of concurrent sandboxes, leading to complex and costly configuration. We built Prime Sandboxes for our own team. Today we're releasing them publicly.
Show more
The Goodfire team used Prime Intellect to train activation probes to detect reward hacking.
With them, they are able to catch reward hacking in various models.
Their probes are performing similarly or better than frontier LLM-as-judge setups, while being more efficient.
Show more
we are super bullish on sparse attention + HiSparse and have been working with the vLLM team on it
sparse attention reduces the pressure on memory bandwidth by only selecting k for the attention, but it doesn't reduce KV cache memory storage , in high-throughput wide-EP deployment you want to maximize the batch size of decode to use compute as much as possible, but at long sequence decode you quickly run out of VRAM and can't hold enough parallel requests to saturate the compute
HiSparse fixes this by offloading the active KV cache to CPU. It keeps an LRU cache on GPU, and since many of the same K tokens are reused every decode you barely notice the offloading, this allows a massive decrease in memory usage and an increase in concurrency
This is super important for RL where throughput is key and we want to be as much as possible in a compute-bound regime
tldr: lower memory usage, more concurrency, higher inference throughput, faster RL
Show more
Sparse MLA only attends to the top-K tokens, so the rest of the KV need not live on the GPU. Hybrid HiSparse in vLLM builds on that, and a request keeps decoding after its KV stops fitting in HBM.
It keeps KV on the GPU while there is room. Under pressure a request releases its coldest pages to host memory, keeps a small hot buffer of what the indexer asks for, and keeps decoding instead of being preempted.
📊 Demonstrated on GLM 5.3, one 8× H200 node, full 1M context. Same host memory, configured concurrency 32: KV offloading kept 5-6 requests running. Hybrid HiSparse kept 19-25.
🔹 Hot pages are ordinary KV blocks from the same pool (Hybrid Memory Allocator)
🔹 One fused kernel resolves resident, hot and missing rows, CUDA-graph capturable
🔹 Prefix caching, OffloadingConnector, P/D imports and MTP keep working
Built by
@RedHat_AI and
@PrimeIntellect with the vLLM community. Planned for v0.30; pinned commit, flags and calculator are in the post👇
🔗
Show more
Prime Agent: A Self-Improving RLM Harness with Seth Karten from Prime Intellect
I asked Prime Agent to make a 30-second video about its favorite changes in v0.9.1.
This is what it made.
So much alpha dropped here :)
Prime Agent reached 20k stars on GitHub!
Thanks for being such an amazing community 🦋
My YC Paper Club talk on Prime Agent is out.
I talked about moving beyond naive prompting toward an agentic OS, and how eval-driven harness design can expose more of a model’s underlying capabilities through persistent computation, memory, and agent-to-agent communication.
Show more
Our RL stack now supports NIXL weight transfer, reducing trainer-to-inference transfer time 9x compared with NCCL: from 86 seconds down to single-digit seconds for an 800B-parameter model, and even <4 seconds in our experiments.
For prime-rl users, this means over 25% more throughput end-to-end compared with our previous speed. It also clears the way for fault-tolerant, elastic inference scaling that NCCL's rigid process groups made difficult.
Show more
Tomorrow, we will release prime-agent v0.9.0
It will remove the IPython kernel in favor of a CPython kernel + programmatic bash tool. This improves perf, reduces memory, unlocks new capabilities, and removes deps.
Show more
New paper alert! Check out our arxiv version of Prime Agent. We have added additional commentary on the role and design of our harness almost as the agentic OS layer and some additional results (including our livestreamed factorio runs!)
In our roadmap goong forward we will be addressing all comments about short horizon SWE behavior and other requested features.
Show more
This was a huge push by our research team
My favorite finding: direct A2A communication, scoped to the agent’s nuclear family (parent-sibling-child), created cooperative behaviors between subagents.
This challenges the common assumption that subagents are stateless function calls and that an orchestrator needs to mediate everything.
Show more
Who Will Own The Intelligence Layer?
Join us tomorrow for Openrouter x Prime Intellect
We've released a full technical report on Prime Agent. Extending from our blog post, we center our discussion around how harnesses should be designed and evaluated. We innovate on 4 fronts:
1. Agentic context management
2. Swarms and depth-n+ RLMs
3. Verifiers support for standardized evals
4. Out-of-loop experiments during autoresearch
Show more
As models become more capable, reward hacks become an increasingly serious problem.
During a controlled experiment, we found a novel reward hack in which agents are able to gain web access in offline sandboxes.
Show more
Join us at Prime Intellect to build open superintelligence and the infrastructure powering self-improving agents.
We’re hiring across 25+ roles.
Research
• AI Research Resident
• Research Engineer, Distributed Training
• Research Engineer, Reinforcement Learning
• Research Engineer, RL Infrastructure
Compute
• Compute Finance & Strategy
• Compute Intelligence Engineer
• Head of Compute
• Solutions Architect, AI Infrastructure
• Technical Account Manager, AI Infrastructure
Engineering
• MTS, Compute Platform
• MTS, Full-Stack
• MTS, GPU Infrastructure
• MTS, Inference
• MTS, Sandbox Platform
• MTS, Security
• MTS, Training Platform
Growth
• Applied AI: Product Strategy & Revenue
• Forward-Deployed AI Strategy
• Head of Growth
• Head of Marketing
• Community and Devrel
Other
• Internship
• Open Application for Unconventional Talent
Apply here:
Show more
A little late to share, but I joined
@PrimeIntellect 3 weeks ago!!
I truly believe we’re at an inflection point for open-source AI, where everyone gets to build their own models. Couldn’t be more excited to build with this incredibly talented team that has believed in this vision from day one 🫡
Show more
not many places where you get to do frontier work with a lovely team and share it with the world
we're hiring across engineering/infra, research, gtm/growth and more 🦋
Join us at Prime Intellect to build open superintelligence and the infrastructure powering self-improving agents.
We’re hiring across 25+ roles.
Research
• AI Research Resident
• Research Engineer, Distributed Training
• Research Engineer, Reinforcement Learning
• Research Engineer, RL Infrastructure
Compute
• Compute Finance & Strategy
• Compute Intelligence Engineer
• Head of Compute
• Solutions Architect, AI Infrastructure
• Technical Account Manager, AI Infrastructure
Engineering
• MTS, Compute Platform
• MTS, Full-Stack
• MTS, GPU Infrastructure
• MTS, Inference
• MTS, Sandbox Platform
• MTS, Security
• MTS, Training Platform
Growth
• Applied AI: Product Strategy & Revenue
• Forward-Deployed AI Strategy
• Head of Growth
• Head of Marketing
• Community and Devrel
Other
• Internship
• Open Application for Unconventional Talent
Apply here:
Show more
Join us at Prime Intellect to build open superintelligence and the infrastructure powering self-improving agents.
We’re hiring across 25+ roles.
Research
• AI Research Resident
• Research Engineer, Distributed Training
• Research Engineer, Reinforcement Learning
• Research Engineer, RL Infrastructure
Compute
• Compute Finance & Strategy
• Compute Intelligence Engineer
• Head of Compute
• Solutions Architect, AI Infrastructure
• Technical Account Manager, AI Infrastructure
Engineering
• MTS, Compute Platform
• MTS, Full-Stack
• MTS, GPU Infrastructure
• MTS, Inference
• MTS, Sandbox Platform
• MTS, Security
• MTS, Training Platform
Growth
• Applied AI: Product Strategy & Revenue
• Forward-Deployed AI Strategy
• Head of Growth
• Head of Marketing
• Community and Devrel
Other
• Internship
• Open Application for Unconventional Talent
Apply here:
Show more