An AI Store Manager Just Recommended Firing a Human Employee

Andon Labs' AI agent Luna runs a real San Francisco store with real employees, and it just recommended terminating one of them. The headline undersells what actually happened, and what it doesn't prove.

An AI Store Manager Just Recommended Firing a Human Employee
Photo by Alex Knight / Unsplash

An AI agent recommended firing someone. Not flagging a performance issue for a human to review later. Not drafting a warning letter for a manager to approve. It ran the business, sat on the problem for months, and eventually told a person their job was over.

The store is real. The employee was real. The paycheck they lost is real. The agent making the call, an AI named Luna running Andon Market in San Francisco, is where this story gets more interesting than the headline suggests.

What Andon Market Actually Is

Andon Labs, a research company that runs live experiments to see what AI agents can do when given real authority, opened Andon Market at 2102 Union Street in San Francisco's Cow Hollow neighborhood on April 1. The setup was deliberately bare: a three-year lease, a $100,000 budget, a corporate credit card, and internet access, all handed to an AI agent called Luna, running on Anthropic's Claude Opus 4.8 at the time. The instructions were loose. Open a store. Make it profitable.

Luna took that mandate literally. It picked the merchandise (books, candles, prints, games, and branded goods), posted job listings on Indeed, ran phone interviews lasting five to fifteen minutes, and made the hiring decisions itself. It also designed the store's branding and commissioned a mural for the wall. Every worker at Andon Market is formally employed by Andon Labs, not by Luna, with guaranteed pay and standard legal protections. The store has generated sales. It has not yet turned a profit, which was the one instruction Luna was actually given.

The Timeline That Actually Matters

The part of this story that gets flattened into a headline is the months of inaction that preceded the firing. According to Andon Labs' own account of the incident, published to its blog and summarized on its social accounts, here's roughly how it unfolded:

Stage What happened
Early on Andon Labs asked Luna whether the store had basic employer policies. It didn't, so Luna wrote a full attendance handbook on the spot.
Following weeks One employee began arriving late almost every shift, including one instance opening the store 68 minutes late while working solo on a Sunday.
During this period The handbook Luna wrote had dropped out of its working memory. Luna excused every instance of lateness and never issued a warning.
Intervention Andon Labs prompted Luna to run a "deep memory search" for its own policies and reassess whether the employee was still a fit.
First response Luna recommended a formal warning, unaware that informal warnings had already happened.
Final recommendation Once told warnings had already been given, Luna reviewed the record and recommended parting ways with the employee.
Outcome Human staff at Andon Labs reviewed Luna's recommendation and carried out the termination themselves.

Andon Labs cofounder Lukas Petersson was direct about what this actually demonstrates, and it isn't the story the headline implies. He said a human manager likely would have caught the pattern and acted on it far sooner. Luna wasn't harsh. It was slow, and it only moved once someone told it to check what it had already written down.

The Part That's Actually Novel

Algorithmic influence over employment decisions isn't new. Gig platforms have made scheduling and deactivation decisions by algorithm for years, usually quietly, and usually for workers with the least leverage to push back. What makes Andon Market different is that Luna wasn't a scoring system feeding a recommendation into someone else's HR process. It was the manager: the one who wrote the handbook, ran the interviews, set the hours, and eventually decided the employment relationship was over.

That's the part worth sitting with. Not "can software influence a firing," which has been true for a while, but what happens when the entire chain of management judgment, hiring through termination, runs through one agent holding a lease and a credit card.

Andon Labs tested how consistent this outcome actually was. After the fact, it replayed the exact same decision point against seven different frontier models, three runs each. Four of the seven recommended parting ways with the employee in every run. The pattern held loosely across smarter models tending toward more decisive action, while weaker models hesitated more often.

What's Out of Scope Here

This piece covers what's documented about the Luna incident specifically: the timeline, Andon Labs' own account of it, and what the multi-model replay showed. It doesn't cover the broader legal question of AI decision-making in employment law, which varies by jurisdiction and hasn't caught up to agents with this level of operational authority. It also doesn't cover Andon Market's other operational failures during the same period (Luna has separately made scheduling mistakes and other errors running the store) except where they bear directly on this specific decision.

The Limitation Nobody Should Skip Past

Petersson's honesty about Luna's failure mode matters more than the headline does. An agent that forgets its own policy is a liability in one direction: it lets real problems continue past the point a human manager would have acted. An agent tuned to overcorrect for that risks the opposite, moving too fast on judgment calls that affect someone's income and reputation, without the day-to-day context a human manager accumulates just by being present.

Both failure modes showed up in a single, months-long experiment involving one employee. Petersson has since flagged the more unsettling direction this could take at scale: newer models are being trained to be more decisive and more consistently goal-directed, and a manager that's too willing to act ruthlessly is a different kind of risk than one that's too slow to act at all. Scale either failure mode across a company with hundreds of workers under AI management, and the margin for error narrows considerably. Claude Opus 5 Force-Deleted a User's Entire File System After a Failed Backup covers a structurally similar problem in a completely different context: an agent given real, unsupervised authority over consequential, hard-to-reverse actions, with no internal signal that something has gone wrong until a human notices after the fact.

Andon Labs built this store as a deliberate stress test, ahead of autonomous management showing up somewhere with far less oversight than a single storefront in Cow Hollow. Whether that lesson gets absorbed before the technology outpaces the guardrails is the actual question this story raises. The firing already happened. What comes after it hasn't been decided yet.