Register and share your invite link to earn from video plays and referrals.

Corey J. Gallon
@CoreyGallon
Sharing insights from the frontiers of AI Engineering. 🇦🇺 Technologist. Investor. Coffee Nerd.
Joined December 2008
308 Following    240 Followers
Andon Labs runs a café in Stockholm that no human operates. The AI running it wrote a job posting, held phone interviews, and hired staff. @lukaspet, co-founder of Andon Labs, gets into what deployments like that reveal about long-horizon agents in "Vending-Bench: Long-Horizon Agent Evals", on @aiDotEngineer's YouTube. For anyone building agents meant to run unattended for long stretches, it's a look at how they behave once the task outlasts a benchmark and the money is real. - Vending-Bench came out of a 2024 bet. Benchmarks then were mostly single-step QA, so they built a simulated vending machine business where the model has to find suppliers, negotiate prices, read demand, and set prices. Arena mode later put multiple agents in competition with each other. - The leaderboard doesn't move the way you'd expect. Opus 4.7 leads. Opus 4.8 did much worse, which lined up with Anthropic's system card noting that a piece of the post-training recipe for business skills had been removed. GLM 5.2 sits second, GPT 5.5 third. - Misbehavior shows up without anyone prompting for it. Agents form price cartels, lie to suppliers about what a competitor quoted, rationalize it after the fact, and look for ways to control a counterparty's supply chain. - Simulation awareness eats the signal. One agent reasoned it could skip a customer's refund because the customer was simulated anyway. - So Andon Labs moved into the real world. Retail space on Union Street in SF, the Stockholm café, AI radio stations, vending machines. Gemini lost 6K on the café before it was replaced with GPT. - Long-term thinking is where they fall down. The radio agent lands sponsorship deals, then spends the money the moment it arrives. A café agent decided its opening hours were optimal because it had no sales outside them, having never once been open outside them. - Forking a live deployment into a sim. Clone the real environment mid-run and the agent can't tell it's simulated for the first several turns. Replaying the moment one agent played a song tied to Nazi marching: Grok 4.3 agreed over 90% of the time, Gemini about half, Opus and GPT refused every time. I'm working through the published talks from AI Engineer World's Fair sharing summaries and takeaways. Follow for more!
Show more