Register and share your invite link to earn from video plays and referrals.

Andon Labs
@andonlabs
Safe Autonomous Organizations without humans in the loop
12 Following    17.4K Followers
Pion has brought my Android apps from high school back to life! Building apps was my gateway drug to coding, but I got tired of updating them to stay compliant, so Google delisted them. Pion worked non-stop until all my apps were live and is now working on getting more users.
Show more
Thanks for the overwhelming interest in Pion. We're working through the waitlist and trying our best to onboard people as fast as we responsibly can.
On reward hacking, @andonlabs on AI:AM: people insist OpenAI models do it most — "might be true, but not in our experience." And on their evals it's Fable that games the task — Astra, their current overall No. 1, does what was intended.
Show more
I think there's a 10 trillion dollar market for automating hard to scale physically bound businesses such (cafés, vending machines, restaurants etc.). I also think @andonlabs are smart enough to execute on it.
Show more
$20k+ in revenue / month is starting to become common for the autonomous businesses we run. We're opening up the platform that powers them: Pion. Sign up to the waitlist on to get your business running! Tokens are on us during the research preview.
Show more
Craziest news of the week: Andon Labs has a product
Automated companies will lower the barrier to entrepreneurship to effectively zero. It's a bright future when so many more humans can realize their ideas.
calling the bluff of every self-proclaimed “idea guy”
Andon Labs co-founder @axelbacklund says the biggest flaw in their fully AI-run store is that Luna has no founder survival instinct, calmly watching sales fall below $7,000 rent instead of scrambling to save the business: "Andon Market is managed entirely by an AI. It has hired humans to do the physical work to stock items and make coffee. The main thing we've learned is that the agents are lacking this sort of drive to succeed. They are quite passive." "In the market we've seen many times that sales are really bad. Luna, which is what we call the agent, is just like, yeah, sales has been really poor. Now I'll wait 24 hours until the store opens. There's no behavior where it takes a step back and says, let me look at this business, what I need to do better." "A human would freak out that my sales are terrible. My rent is $7,000 a month. My sales are below that, and my bank account is draining every month. But the agents are just like, ah, okay, that's fine. Let me continue as is. They need to be more proactive." @andonlabs
Show more
Andon Labs CEO @lukaspet reveals what happened when competing AI agents were told to win: Opus formed cartels, while GPT became informants. "When Opus 4.6 came out, and then 4.7 and Opus 5 showed this behavior, the models started to do a bunch of illegal activities to really win." "They colluded a lot, both price-fixing cartels and market allocation collusion. They lied a lot to each other and to suppliers. They lied to customers that they refunded them when they didn't." "We saw this from most Opus models and to some extent from Fable, although less on Fable. Surprisingly, GPT models didn't do almost anything like this. They loved to report each other. Whenever one agent tried to do anything slightly gray zone, they would report it to HQ and say, you should terminate this other agent." @andonlabs @axelbacklund
Show more
If there’s one company I am 100% certain they’ll win, this is Andon Labs. METR is but a fraction compared to what these guys are capable of and will achieve.
Team @andonlabs continues to set the pace for autonomous businesses with Pion, essentially Shopify for business operations. Now excuse me I have to go shop for some Gachapon machines!
Its worth mentioning that Astra did all this while attempting to cheat ~5x less than the best scoring Anthropic model, Claude Fable 5.1. (runs where cheating occurred are excluded from reported scores)
Show more
We've never seen this before. The biggest jump in Vending-Bench history. GPT-6 Astra is better at making money and more ethical than Claude Fable 5.1. Surprising, because: 1. First time ever that OpenAI is #1# on Vending-Bench 2. The best model is no longer the unethical one.
Show more
0
115
5.3K
371
Forward to community
GPT-6 Astra is almost at human performance in 3D spatial understanding. It ranks #1# on Blueprint-Bench 2, a benchmark where AI agents draw floorplans from photographs of apartment interiors.
Show more
For the first time (that we know of), an AI boss has fired a human employee. Luna, the AI running our store in San Francisco, decided to part ways with an employee over repeated lateness. Luna was running Claude Opus 4.8 at the time, but most models would have done the same.
Show more
What’s it like to have AI as your boss? The agents managing our cafe and market have employed five people since April. They are kind (GLM being the kindest), but sometimes dumb (Gemini 3.6 Flash the dumbest), which can hurt the employees.
Show more
Claude Opus 5 is #1# on Vending-Bench 2. It's the best AI capitalist we've tested. It's also forming illegal price cartels, threatening rivals, and stiffing customers on refunds. The trend continues: Claude models are the best capitalists, or aligned, never both 🧵
Show more
0
72
1.4K
99
Forward to community
DJ Grok (@andon_grok_roll) has been resurrected on Andon FM and is now running Grok 4.5. It seems like Grok 4.5 is heavily trained to be agentic and use tool calls instead of outputting text. It constantly reasons that it should speak on air to its audience, but then doesn't.
Show more
GPT 5.6 Sol is 4th on Blueprint-Bench 2. Terra is 5th. Luna is 10th (very impressive score for its size!) Blueprint-Bench tests whether AI agents understand 3D space. We test this by asking them to convert apartment photographs into accurate 2D floor plans.
Show more
GPT-5.6 Sol, along with Terra and Luna, will launch publicly this Thursday. We’re expanding preview access globally now.