Frontier models are simply too expensive and slow for the majority of use cases, so we see models like SWE-2 becoming the daily driver for most.
Training trillion-parameter coding agents at scale isn't easy though: typically, each step launches thousands of rollouts, each with its own isolated environment.
Cognition uses Modal's sandbox infrastructure for the rollouts behind SWE-2. Congrats on the launch!
More on how we scale Sandboxes:
顯示更多
Introducing SWE-2, our closest model yet to the frontier.
On leading evals, it scores on par with recent frontier models – at up to 70% lower cost.
We scaled RL to multiple trillions of parameters, with a refined recipe that pushes the Pareto curve on both capabilities & cost.
顯示更多