Register and share your invite link to earn from video plays and referrals.

Thorsten Ball
@thorstenball
Working on @ampcode. Writing Once authored and
1.1K Following    53.7K Followers
Hot Take Thursday!
Would you prefer to save money with dial-up when your competitors are using fiber? Recorded on Hot Take Thursday, @sqs and @thorstenball are back with a new Raising An Agent. In this episode they discuss how cheaper models end up costing more, why they do not not wait for agents, running evals in Orbs, and the hot takes Thorsten left out of his X post last week. 00:00 Intro: Engineers who forget the history of technology 01:55 Hot Take Thursday & a €50-a-week token budget 05:14 Why cheap models cost more than frontier models 08:11 Token budgets & who gets priced out 14:06 Disrupt yourself with small teams & unlimited tokens 16:38 Agents go horizontal across the software lifecycle 20:56 Rollouts, bug reports & "I'm feeling Pucky" 23:53 Code review, monorepos & other walls around the agent 25:30 How to show a cautious company what agents can do 29:48 Why most people haven't seen what agents can do 31:57 Don't wait for your agent 33:38 Simple evals & agents testing agents in Orbs 38:27 Why Amp has runners too 40:47 Follow your energy & stay aligned 46:33 The hot take Thorsten left out 54:14 Do languages & frameworks still matter?
Show more
the moat is literally just caring about your work and being proud of your output and holy shit it does not seem to apply to 99% of people
0
109
7.5K
513
Forward to community
My maybe controversial opinion: If TTFT is important to you, you're doing something wrong. Don't wait for the agents, make them wait for you. Don't stare at a single transcript. Do something else. (Yes, there's a tiny number of exceptions to this.)
Show more
playtesting every harness again, i cannot believe how you folks don't lose it over the TUI taking over 50ms to start, and TTFT >2s - these companies are achieving AGI and here we're stuck with the slowpoke harnesses wtf
Show more
Launched just now in amp - if you have the mac app installed, it can automatically spawn runners for you! No more pesky terminal windows lying around :^)
Increasingly convinced: people who do best with agents are those that are good at shipping things in small pieces. Someone who could ship a feature in 4 commits going to main now does better with agents than the guy who always ended up with a 6k PR on Friday afternoon.
Show more
Another thought I had yesterday in that conversation: These models are so much more now than text-to-code converters. Think bigger! Aim higher! Now that the models can contribute across your whole stack, vertically, it's to let them loose horizontally: Let them monitor deployments, let them debug prod, let them do ops, let them help end to end. If you think "they'll screw this up" that's on you and your codebase. Either you need to spend more tokens and build up instincts or you need to make your codebase and company processes friendlier to agents.
Show more
Unexpectedly found myself in a conversation about agents yesterday evening with someone who now has to use Qwen at work, because company wants to save money. Really struggled with expressing how much of a category change using latest frontier models is vs. using Qwen.
Show more
Unexpectedly found myself in a conversation about agents yesterday evening with someone who now has to use Qwen at work, because company wants to save money. Really struggled with expressing how much of a category change using latest frontier models is vs. using Qwen.
Show more
When you have friends around the world you realize: It's always orbin' time somewhere
still can't get over how fun (and productive) multiplayer amp orbs are frontend changes used to be sooo slow to iterate on
"the best harness in the market"
@AmpCode is the best harness in the market, there's no point even doing a tier list
Same prompt, four modes/models: - GPT-6 Sol - Opus 5.5 - High (GPT-6 Astra) - Ultra (Fable 5.1) Here's the prompt: "Look at the last two news posts we published on runner functionality: one runner is now enough, and one runner many worktrees. I now want you to write a news post that's similar but for the functionality behind the `shared-runners` feature flag. It should probably also contain a screenshot of the picker that shows two normal runners and two shared runners, like `gpu-runner` or `macos-builder` or something and some nic elooking avatars and good data. And it should succinctly explain how it works and how people can use it." Costs: - Opus 5.5: $9.54 - GPT-6 Sol: $3.10 (estimated list price, using sub) - High: $7.58 (estimated list price, using sub) - Ultra: $15 Notes: - Ultra took by far the longest, Ultra is the only one that didn't create docs (I didn't say they should create docs but I think the other models looked at previous commits and figured out they should do that?) - Ultra was the only model that didn't prominently display the "give me the grant to upload assets at the end" - I think Opus 5.5 is the best result. It really feels like more creative (headline) and it nailed the screenshots. - High is my second favorite.
Show more
I'll say this: I've gotten more out of using orbs and asking agents "test this e2e & give me irrefutable proof it works" than out of 80% of tests I've written and ran over the last 15 years. It's truly remarkable that we can now ask machines: "Did you manually test it?"
Show more
I agree with @thorstenball, I think unit tests are dead in the water. The ones the models write are terrible, at best just doubling total LOC. Inverting the testing approach - heavy e2e/black box/golden master, reaching for lower levels only if necessary, works better for me
Show more
"this technology will create little moments of delight, whimsy, and value, in everything we touch" yes!
people don’t seem to get the idea that this technology will create little moments of delight, whimsy, and value, in everything we touch. the whole reason you don’t fine tune a classifier for one task is so users can shape the software to themselves, not vice versa
Show more
I know people who spent $12k in tokens and built nothing. @Westpac spent $12k in tokens on @AmpCode and built a new intranet for their 35,000 employees (vs. ~$3–5M pre-AI cost). "Amp is unlocking a huge amount of value in engineering" from
Show more
"Didn't you say intelligence is going to get cheaper?!"
Quite a few people clowning on me for saying that UI will be generated on the fly. Two things: 1. You don't have to generate it all from scratch, but dynamically assemble widgets. 2. Tried chat jimmy? In a few years, we could have today's frontier intelligence at this speed
Show more
Most predictions I see are still way too conservative. Here's mine
What's the % of developers working on pace makers? 0.05%? Less? What comes up in every discussion about software processes and that maaaybe some can be changed? Pace makers.
Runners can now create Git worktrees, one per thread. No more threads fighting over a checkout.
It's orbin' time
Search to Siri animation in macOS 27.2 beta 2