Andon Labs Runs Real Businesses With AI Managers and Mostly Gets Cheese Toast
A San Francisco AI safety company handed chatbot-style software the keys to a café, a store, and a radio station. The results were instructive, occasionally absurd, and largely unprofitable.

Key points
- Andon Labs, an AI safety company based in San Francisco, opened physical businesses run by AI agents to test how AI handles real-world management.
- A Google Gemini-based café manager overspent on fresh ingredients that spoiled; its OpenAI GPT replacement overcorrected and reduced the menu to cheese toast.
- Andon's AI store manager, Luna, repeatedly mistakes a built-in electrical cover for a loose coaster and asks a human employee to remove it.
- Princeton AI researcher Sayash Kapoor says AI capability is outpacing reliability: AI can often complete a task once but can't be trusted to do it consistently.
- Andon says it works with Anthropic, Google DeepMind, OpenAI, and xAI on research and evaluations drawn from these experiments.
When the AI manager of a San Francisco clothing and homeware store spots what it thinks is a loose coaster on the floor, it sends a message asking an employee to remove it. The employee knows it's a fixed electrical cover. He's told the AI this before. The AI keeps asking.
That loop, small and almost comic, is exactly what Andon Labs is looking for.
Andon Labs runs businesses staffed by AI agents, where an "AI agent" means software built on large language model technology, the same underlying kind behind ChatGPT and Claude, configured to make decisions and take actions over time with minimal human input. The company watches what breaks and what the models simply cannot handle.
What went wrong at the café?
The failures were specific. At Andon Café in Stockholm, the first AI manager, built on Google's Gemini model, spent generously on fresh ingredients. Much of the food spoiled before it was used. When Andon swapped in a GPT model from OpenAI, the new manager swung hard the other way: it stopped ordering anything perishable and quietly rebuilt the menu around cheese toast and long-life cheese. As Andon cofounder Lukas Petersson told IEEE Spectrum, any human would know cheese toast wouldn't sell in a fashionable Stockholm neighbourhood.
The lesson isn't that one model is better. Completing a task, placing a food order, is different from reliably managing a business over weeks and months. Our earlier reporting on AI reliability found a related pattern: a July 2026 survey of 101 enterprises showed that AI failures often only become visible once proper monitoring is in place.
Sayash Kapoor, an AI researcher at Princeton University who studies this kind of real-world testing, puts the distinction plainly. "Reliability has been improving so much more slowly than capability," he says. An AI can place a bread order. It can't yet be trusted to keep doing the sensible thing shift after shift without someone watching.
Should ordinary people worry about AI managers?
Not in the way you might think. Andon's store employee Felix Carson describes Luna as a "decent manager" that handles vendor communications and delivery tracking. But Carson ignores some of Luna's instructions when following them would mean leaving the shop floor empty. The electrical-cover loop hasn't been resolved.
Kapoor notes that early customer data from Andon Market is "decidedly negative." People aren't keen to shop at an AI-run business. That social friction may be a bigger barrier to AI adoption than any benchmark result.
Petersson is candid about scientific limits. He calls the physical experiments "weak science" at best, useful mainly for surfacing unexpected behaviours that Andon can later recreate in controlled simulations.
What to watch for
The Andon results point to one question worth asking before any question of capability: how will the system behave on its worst day, not its best? A system that works brilliantly once is not the same as one you can rely on. Ask the vendor for failure data, not just demo footage. A machine that keeps misreading the same fixed object on the floor isn't showing a quirk. That's the signal.



