I spend a lot of time thinking about what has to happen for better models to become better software. The research progress is remarkable, but there’s still a lot of work between a model gaining a capability and someone being able to depend on it.
That work is interesting in its own right. Understanding a codebase, fitting into how a team operates, knowing when a result is wrong, making something reliable enough that people can stop checking every step. Better models help with all of this, but someone still has to take responsibility for the whole system.
It is now clear that it’s nothing short of critical that some of the companies doing that work are model-independent. Different models are useful for different things, and those tradeoffs keep changing. We should be able to choose based on what works for a customer, and change our minds when the evidence changes. Customers shouldn’t have to make a long-term bet on a research lab to improve how they build software. We shouldn’t have to trust the creator of the model to do what’s best for their customers. We should force them to in order to compete.
There’s a business model question here too. If we find a way to do the same work reliably with less inference, that should be a good thing for both us and the customer. If another lab makes a breakthrough, we should be excited to put it to use. Independence doesn’t automatically create that alignment, but it gives us more room to build around it.
I’m interested in where intelligence research goes. I just don’t think every AI company needs reaching AGI to be its organizing purpose. Helping people build things they couldn’t build before is a substantial goal on its own, with useful measures of progress along the way.
That’s how I think about Factory, and what makes today’s funding meaningful to me. More room to work on these problems, with our attention on the people using the software.
Best way to reduce your p(doom) is walk through the causal chain of any AI-driven ex risk scenario.
There is a huge amount of hand waving and sci fi reasoning that happens in the middle of these chains. Great thread from @DavidRBellamy
The companies that win will own the verification, state, learning loops, and workflows around the model. We chat about this, open models, and more.
Thanks for having me @HarryStebbings!
We show how a model-independent AI system can radically improve performance of any frontier model on the hardest problems in software development. New Factory research on long-horizon agents.
Models can win the IMO because a proof checks itself. A codebase doesn't, so agents declare victory early and fall short of human-level work. This makes solving complex software far harder and more interesting than solving math problems.
Like the constraints human organizations impose to ship production software, agents can organize themselves and build new validation strategies to verify the unverifiable.
This is likely the most advanced program reverse engineering system in the world right now.
This is a new era for cyber threats. AI is the single most powerful tool for defensive action. We must act swiftly to harden the digital infrastructure that acts as the fuel of progress.
Proud to support @openai's efforts here!
Proud to have supported AGENTS.md in our product since its creation.
Capturing your customers with proprietary formats in these early days of AI isn’t optimal strategy. Making products people love will create more enduring believers in your vision of the future.
One of the biggest discussions in AI right now is ROI. It's difficult to build a strong relationship between the outputs of agents and their corresponding impact on the organization.
This is a big step in solving this problem at scale across the entire enterprise.