Exo Harness now runs natively on Runta.
In FrontierHarness Eval, Exo was cheapest per completed task at $1.05. On the hardest task it hit its 51-step cap and quit at $1.46 while others kept spending.
That is the harness you let rewrite itself. Exo's Executor holds no durable state, so the agent can modify it. History, artifacts, secrets and sandbox lifecycle sit in the Harness, out of reach.
Runta provides a resumable environment, so the Exo Harness can self-evolve freely. Exo never holds the model API key, only a stub. Our egress gateway injects the real one at the provider. Code the agent wrote can read whatever the Executor can, and all the Executor has is a stub.
Front page of HN, 1700+ posts on X.
Turns out a lot of people have been wondering what the harness layer actually costs , and whether the expensive ones are any better.
The harness war is on.
A year ago the question was which model. Now it's which harness.
Pi, Exo, Claude Code, Codex, DeepSeek Harness and 4 others. Same model, same tasks, same runtime. 360 runs, 2 billion tokens.
Pass rates: 50% to 67%.
Cost per pass: $1.05 to $18.34.
Introducing FrontierHarness Eval. 🧵
Yes, we are hiring!
We are building the execution layer for next billions of agents, and we are looking for high-agency engineers:
Core Runtime (kernel-and-VMM-layer)
Cloud Infra (IaC and multi cloud)
Product Engineering (Agent Experience and DX)
Small team. High talent density. High ownership.
Apply here:
Today we're announcing @runta's $20M seed, led by @a16z.
Software is constrained when you write it. Agents have to be constrained while they run.
Runta is the execution layer that controls what AI agents can actually do.