Register and share your invite link to earn from video plays and referrals.

Walden
@walden_yan
I sometimes code @cognition
727 Following    14.6K Followers
Agents with MacOS VMs are very powerful. Just one example: Devin can test android and iOS apps side-by-side.
Special delivery: Devin just got a Mac 🍎 Now Devin can: 1. Build & test apps on its own Mac VM with iOS simulator 2. Send a screen recording via Slack 3. Send a TestFlight link so you can start using it 📲
Show more
We independently benchmarked Devin Fusion for its release today - this is the first time a multi-model coding agent has been included on the Artificial Analysis Coding Agent Index, and it effectively retains Claude Fable 5.1 and GPT-6 Astra performance while reducing costs Devin Fusion runs a frontier lead model with a cost-efficient sidekick. We tested configurations from Cognition combining frontier models from Anthropic and OpenAI with their new SWE-2 (medium) as a sidekick model. Configured with Claude Fable 5.1 (xhigh) + SWE-2 (medium), Devin Fusion scores 62 on the Coding Agent Index v1.5, while with GPT-6 Astra (xhigh) + SWE-2 (medium) it scores 59. The Fable configuration has the higher score, while the Astra configuration is 43% less expensive and completes tasks 31% faster. Congratulations to @cognition on the release! See below for our results and analysis 🧵
Show more
0
94
1.5K
87
Forward to community
We've seen far more demand for SWE-2 than anticipated, leading to some capacity shortages – thank you for your patience as we scale up compute! We'll be extending the free SWE-2 promo into October to make sure everyone gets a chance to try it.
Show more
A big improvement for CLI fans who want to use more Astra or Fable. We've tuned Fusion so that the delegation is unnoticeable day-to-day. Worth giving it a try. Hats off to the team that made this possible @joon_h_lee @spdling @mapeiyuan @morgantepell @maaslalani
Show more
Introducing Fusion in Devin CLI The most efficient frontier harness for Fable & Astra; 39% cheaper across coding benchmarks. Pick your favorite model for planning and a cost-effective model for execution.
Show more
We launched SWE-2. Its a post-train on Kimi K3 that takes price-performance into account. And it's free in Devin CLI for the next while. While I don't claim it's smarter than Fable or Astra, it's amazing that we can even get close at such low cost.
Show more
SWE-2 is post-trained on Kimi-K3 and proves that our RL recipe continues to scale on stronger base models. On FrontierCode, SWE-2 achieves a score of 50.0%. It beats SWE-1.7, Grok 4.6, and GPT 5.6 Sol while matching Fable 5.1 at 64% lower cost.
Show more
Automated UI testing feels like AGI in Devin 10 examples (running macOS + Linux Cloud Agents) 🧵 What happens (automatic for any update): • Creates programmatic test plan • Boots the app with secrets & credentials (auth, etc.) • Autonomously clicks, types, and checks the UI like a human using a computer would • Records a test • Edits the video, attaches step by step walkthrough of the actions it took • Sends you a concise video artifact As we move away from reading every line of code, new advanced ways of testing become more valuable. 1. Testing a platform for signing documents electronically
Show more
2 years ago, cloud agent didn't work Today, I meet teams running entirely on cloud agents Proud of our team, and still so much more ahead
The world needs far more software than it can build. Cognition exists to change that. We’ve just raised over $2B at a $48B valuation, led by a16z, Accel, Founders Fund, General Catalyst, and Avenir. Since our round in May, run-rate revenue has grown from $492 M to almost $900 M.
Show more
Fable cheaper than Opus -- what a crazy upset Cost improvements across Devin today
Introducing Fable 5.1 in Devin Fable-level intelligence is now 54% cheaper – making it even cheaper than Opus – due to a change in caching. Devin’s Fusion harness is now even smarter and cheaper than before, matching Fable 5.1 on FrontierCode at 47% lower cost. Here's how:
Show more
Minor detail: I spent an embarrassing amount of time trying different harnesses for intelligent range detection and line folding Users don’t even notice it because it’s designed to be hidden. Until they go back to GitHub
Show more
smart diffs on @DevinAI is very very good, im surprised more agent harness companies have not done this as well.
Cool $1M initiative to make AI mathematics progress Excited to be supporting from @cognition
A few @Caltech friends and I are hosting the first math hackathon with frontier labs/startups. We're giving $1M+ of compute to 100 teams to solve the hardest open problems -- competitors range from IMO golds to frontier/neolab folks to math PhDs/profs. DM me if your company wants to be involved :) we are launching soon.
Show more
Come work with us!
Get your slippers on and join us for a walk around our SF headquarters, where we're doing it all with Devin. Even ordering our morning coffee.
There are some sick companies outside Silicon Valley. It’s easy to get lost in the tech bubble. At Cognition, visiting our customers is one of the ways we keep ourselves grounded
In Prague, Czechia, you can order a week's worth of farm-fresh groceries and have it at your door in an exact 15-minute window. Milk in glass bottles from a local family farm. Fish caught yesterday. Strawberries picked that morning. This is all fueled by @rohlikgroup, a powerhouse from Czechia. It’s profitable, it did >$1.3B of revenue last year, and it's accelerating European grocery delivery with Devin.
Show more
Some personal news: I’m joining @cognition as Head of Creative
The world is entering “creative mode”, and teams are only limited by their ambition. My favorite use for Devin isn’t even writing code. Every day, my Devin reads every customer report and escalates all the top feature requests to our eng team. Eyes everywhere.
Show more
At Cognition, one of our values is to go for it all: when faced with a tradeoff, pick the ambition-maximizing direction. When we launched Devin in 2024, we envisioned a future where every team had an infinite army of junior engineers. We were early and Devin wasn't good enough. Now it is! The only constraint left is how ambitious you're willing to be.
Show more
One of the best parts of working at Cognition is getting to work on some extremely sick tech in industries outside of Silicon Valley software
GE built America's first jet engine in 1942. Today GE Aerospace's engine technology powers 3 of 4 commercial flights in the world. Now they’re using Devin to build more, faster. One team nearly doubled its engineering output. Their software optimizes routes, reduces fuel consumption, and monitors anomalies. Their software team's goal is to fly more without adding another airplane, runway, or pilot.
Show more
highly effective Devin automation that costs me under $1.50 per day Every day at 5pm: - Devin trawls through our logs and finds the slowest queries in the app - Creates a linear issue for each - Fixes them This would have taken a $200K per year engineer at least 3-4 hours per day previously.
Show more
Software abundance: a brain dump on the state of the job market and the software industry. My feelings have swung between optimism and pessimism over the past ~8 or so months, and have ultimately landed on the extreme end of optimism / positivity, I'll try to explain why. I've thought about this a lot. In November 2025 if you were a software engineer and tried Opus 4.5, you knew it was basically "over", or you were in denial that it was over. Over in the sense of our identities as programmers no longer was going to mean what it used to mean, typing code etc.. "does this mean less demand for what I do since more people can do it, will I be valued less, will the cost of software go to zero..." In terms of the software job market, a large portion of it is booming. Many companies are being born, exploding in value, revenue, and hiring. People who became AI pilled in their approach to software engineering have done extremely well, those who haven't have not. People who have pivoted into new and fast growing verticals, companies, products enabled by AI have done well, people who stuck with larger teams / companies that either are moving slowly or are being disrupted have had a hard time or are being laid off. People who still distinctly identify as a "frontend / backend / etc.." have had a harder time than people who have said fuck it and 10xed themselves by leaning harder into AI and objectively broadening their skillset + what they bring to the table. It's crazy that I can talk to someone with the former mindset and it's doom and gloom while people in the latter literally are making more than they ever have, along with more opportunities than they have ever had. My craziest realization though is more around the state of the software industry and software abundance. Since joining @cognition I've realized Jevons paradox is real and it's incredible to see. The number of software companies is exploding, the amount of software being shipped is exploding (teams we work with regularly 10-20x their shipped PRs), the number of people starting to write software is exploding. I've also realized that just because code is easier and faster to write, good software is still hard to get right, needs care, and requires good judgement. AI allows you to ship 20x more code, but so can your competition. So while the software industry has always moved quickly, it continues to accelerate with AI. Our roadmaps just get larger, our bar is higher, the quality of the software we ship is better. I was scared that the craft of building software was going away, but I've realized it's actually the opposite. The majority of what I'm seeing automated is the type of work that engineers hate - fixing bugs, doing migrations, upgrading dependencies, repeatable tasks. Agents are *very good* at this type of work. What agents still lack is the judgement, curiosity, context, and taste of a human plugged into to the outside world. So as we automate the work we don't enjoy, we're getting more time to experiment, innovate, and get back to doing the type of work we love the most, why most of us began programming in the first place. I'm at a place where I'm more excited and optimistic than I have been in my ~14 year career.
Show more
0
127
1.6K
159
Forward to community
Using agents 24/7, it feels so obvious why certain SaaS (observability, ticketing, CI, cloud) are benefitting massively from AI Must be scary to being software investor who’s never used agents or these products
Show more
Huge Atlassian quarterly beat. There was a misplaced thesis over the past 6 months that somehow agents would be bad for certain software categories. There’s definitely truth in this in some areas, but many were parsing this poorly. In a world where agents are generating 100X more code, processing massive amounts of data, or making decisions across your systems, the role of the platforms that manage this data and these workflows becomes more important, not less. Enterprises care about governance, security, compliance, guardrails, safe access to data, and many more critical capabilities that go into these systems of record.
Show more
My take 24 hours after Fable 5: Your organization will likely not scale with the exponential curve of AI. I'l just come out to say: This should be a wakeup call for engineering teams. Set up your cloud software factories. Now. Models can now fix impossible bugs, UI-test the hardest flows, writing extremely good code, etc. I have't opened Datadog manually as far as I can remember. AI should be the first-line defense for bugs and feedback. Humans should only look at PRs after an AI has already reviewed it. AI should generate screen recordings of any PR before a human eye even reaches it. The agent should just prompt itself most of the time. Ex. (pictured) our ui feedback channel manages itself, creates tickets, assigns itself automatically You might also be worried about cost. Anthropic, OpenAI, and other labs will likely continue to put out bigger and more expensive models. But, we will also continue to get more capable small models. Not everything will need the smartest models. It's about having the organizational harness in place to continue taking advantage of this rising tide. Moreover, if you use Devin, we've already optimized our harness a bit, and Fable is actually only ~40% more expensive in practice (vs the 2x people assume). I'm honestly pleasantly surprised - it might be higher ROI than you think. Anyway, if you take anything away, engineers shouldn't be manually picking up tickets, humans shouldn't be digging into logs themselves, rethink what you do with your time that shouldn't just be an AI. We need to rethink what humans spend their time going.
Show more
One think I really like with Devin Auto-Triage vs most home-grown SRE/bugfix automations: It’s structured as a manager agent + a subagent fleet The system has the full context + running memory. So I can ask for reports, to dedup bugs, to remember to tag certain people, and more
Show more