Register and share your invite link to earn from video plays and referrals.

Abhay Singhal
@_AbhaySinghal
226 Following    4.3K Followers
I made Claude compete against itself. The smartest model does not automatically make the best agent. To prove it, I took the same Opus 4.6 model, initial prompt, and empty repo and but it inside two different coding harnesses: Claude Code vs. @FactoryAI Droid. The task: clone Excalidraw from scratch, inspect the original in a browser, implement its key interactions, and verify the result. Factory Droid: 8 minutes, 20 tool calls, $1.60 Claude Code: 19 minutes, 40 tool calls, $1.87 The only difference was the harness. Get Free Factory credits here:
Show more
Model routing can happen across layers, each with a different set of information. The final choice belongs in the harness because it can: - Reduce cost while maintaining performance, accounting for the prompt cache - Shape jobs and choose their models together - Learn continuously from task outcomes More on why, plus data and learnings from building @Factory Router
Show more
You can only optimise what you can measure. Time to start ROI-maxxing
Introducing Agent Effectiveness. Measure Factory usage across cycle time, work priorities, and shipped artifacts, so you can actually see what your AI spend produces.
10,000 human-days of modernization work done in 45 days. 45 day database migration completed in 6 hours. See how @Comarch is building its own software factory.
This is why @FactoryAI is hiring EPD in SF only. We’re also hiring fewer engineers than before, even as growth continues, so there’s less need to hire in more markets. A top engineer can amplify their output by orders of magnitude with agents, whereas others would just multiply bad decisions and bad code faster. Many of our largest projects are already bottlenecked on a single engineer because the shipping pace is so fast. Adding more people often wouldn’t meaningfully speed up timelines.
Show more
We’re seeing model routing deliver significant savings across large enterprises and all workflows. Full model independence and reliable harness-level routing enables this access to the Pareto frontier. Excited to share this work with everyone! More to come. Expect more savings as efficient models and routing continue to improve
Show more
Router has been in production for a few weeks and the savings have exceeded the original benchmarks. Measuring billed usage vs the same workload priced at Opus rates, we’ve found: 43% aggregate cost savings 61% of sessions are at least 80% cheaper
Show more
GLM 5.2 Fast is now available in Droid, hosted by @FireworksAI_HQ. GLM 5.2 is one of the most popular models in Droid. Now you can build with the same quality at much higher speed.
Show more
Introducing Droid Shield 2.0: learned secret detection for safer autonomous engineering at scale.
Today, we're announcing Factory 2.0: from coding agents to software factories.
0
144
1.5K
148
Forward to community
We're in the midst of another phase shift from agents to factories as focus moves from busyness (token usage) to outcomes and ROI across organizations. Deploying agents to tokenmaxx and then setting individual token budgets may improve productivity, but it does not necessarily improve outcomes at scale. 100x'ing an enterprise is not the same as 100x'ing an individual. Coordinated agentic systems with efficient resource allocation and clear measurement of outcomes are needed to unlock the full value of AI. Underlying this is a set of closed loops so the system continuously learns and self-optimizes. Human context and governance will only be more crucial as organizations build their software factory. Good judgement and honing one's craft are becoming even more valuable as the way we work transforms.
Show more
This is how we have more 9s than any single provider. In an era where status pages look like festive lights, it's not enough to pick the right model for each task. Every model needs to finish the work reliably.
Show more
Reliability is built into @droid. Factory Router has auto-failover across model providers, so your sessions keep running even when one of the providers goes down. We also provide dedicated compute capacity for enterprises, and we reserve a guaranteed TPM (token per minute) allocation. We believe that diversifying LLMs in your agent stack is the future, and the way to achieve the highest reliability in production.
Show more
Model routing is an important thing Controversial idea: the frontier labs will want their AI harness to be the moat, but ultimately the best case for consumers is that model capabilities flatten and commodify Preview of the AI Harness Wars of 2027
Show more
0
138
1.3K
77
Forward to community
Automatic behind the scene routing in user interfaces (instead of model picker) will redistribute value capture and usage towards many more models than just frontier ones (especially towards open-source/smaller/cheaper ones). Because it removes the cognitive load for the final user to have to switch models (which is too high leading to defaulting to frontier for all calls with model picker) cc @antonosika @bgurley
Show more
What needed a frontier model last year runs on a more efficient one today, yet AI costs keep climbing. A higher token bill doesn't mean more work is getting done. Engineers default to the most powerful model out of fear of losing performance, so routine work runs the same premium path as the work that genuinely needs the expensive model. Factory Router cuts token spend by 20-25% while maintaining frontier performance. It automatically selects the right model for each task, and routes across providers for reliability. Designed for agents, it switches models only when the gain is worth rebuilding the prompt cache. We mapped the cost/performance Pareto frontier across benchmarks. Near the top it's nearly flat: cost drops sharply while performance barely moves. Then it bends hard: the most aggressive routing we measured cut Terminal-Bench 2 to 56% of Opus cost but dropped pass rate to 81%. Factory Router operates on the flat stretch, right before the bend. As efficient models improve, more work crosses into what they handle just as well. With automatic routing, users will see growing savings at frontier performance.
Show more
Introducing model routing to Factory. Factory Router picks the right model for every task, automatically. Maintain frontier performance while cutting costs by 25%.
Today we're opening access to Droid Computers: persistent machines for remotely orchestrating Droids. Spin one up in Factory's cloud or turn your machine into a Droid Computer. Either way, Droids have a dev environment with its own filesystem, credentials, and configurations.
Show more
Claude Opus 4.7 has arrived in Droid. Available 50% off for all users through April 30th.
Today we're releasing the Factory desktop app. A native interface for autonomous AI agents that work across every part of your software business.
0
110
942
72
Forward to community