It's time to apply a true engineering mindset to deploying coding agents. There’s too much hand-waving around what agents are best, which models to use, and how to optimize ROI from coding agents over time. The solution is to set up a closed-loop system in the cloud where all of your agents are tracked and measured against your own data and workflows, so you can adjust your setup based on actual data and not vibes.
The emerging category of infrastructure that supports this approach is the cloud software factory. Software factories are automation loops around the SDLC, comprised of agents that triage, spec, implement, verify, review, monitor, etc. Done right, these factories allow for measurement, improvement and automation over time, and can help prove that you are doing agentic engineering the right way.
Factories should follow these design principles:
- Factories should be defined as code, version-controlled, and have their definitions editable by agents.
- Factories must live in the cloud to support team access, central data storage, and automations.
- Factories should have a runtime that is API driven, not UI first.
- Factories should come with built-in evals, improvement loops and benchmarks so you can ensure improvement over time.
- Factories should be inherently multi-model and multi-agent, so they can take advantage of improvements in models and harnesses.
Your goal is getting to a “closed-loop” factory: one where all of the data, observability and improvement features are baked in, so that agents (and humans) can use that data to improve the factory over time.
That last point on data collection and self-improvement is especially important to get right. If you’ve set up your factory properly, these agents take their learnings and propose changes to how the factory operates. They do this by creating diffs against the factory definition, which – assuming your factory is defined as code – is easy for them to change.
The high-level flow for self-improvement comes in three steps:
1. Factory agents do triage, implement, verify, etc. – they build the product.
2. Scorer agents periodically grade that work along dimensions that you define, like cost, quality, verbosity, etc.
3. Self-improvement agents review scores and suggest improvements. Humans review those suggestions as PRs on the factory definition and merge improvements.
The way to think about the factory approach is as meta-engineering.
You should invest now in infrastructure that lets you measure, test and automatically improve the SDLC. The longer you wait, the more catch-up you will have to do and the more tokens you’ll burn in the meantime. Building this infra is possible, but a big endeavor, so as you evaluate factory platforms you should look for ones that have the right primitives to help you scale software development today.
If you're team is building a software factory, let's get in touch:
Show more