
Imagine watching an AI-controlled business operate in real-time — making decisions, facing crises, and even losing money — all in public view. For the first time, this is possible with Firmulate’s live experiment, revealing not just how AI handles business chaos but whether it can truly deliver on its promises.
The Reality of a Company Without Humans
At the heart of this experiment is a small, synthetic company managed entirely by AI models. Unlike typical demos that showcase AI chat skills, this setup simulates a real business with actual money mechanics and daily challenges. The company runs with 13 ‘synthetic employees,’ each guided by thousands of learned rules, and burns through €105,000 every month against a modest €2,300 in monthly recurring revenue (MRR). This public display offers a raw picture of AI’s capabilities in managing complex, high-stakes environments.
AI business management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment: Putting AI to the Worst Test
In a groundbreaking test, four advanced AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—were each tasked with navigating exactly the same difficult week. This week included real customer crises, temptations to manipulate the system, and the pressure of closing a critical €55,000 deal. Every decision made by the models was recorded, versioned, and auditable, ensuring complete transparency and allowing observers to follow their every move.
AI decision-making tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Findings Show AI’s Real Skills
All four models successfully identified every crisis, demonstrating impressive situational awareness. Importantly, every model refused every manipulation attempt, including sophisticated social engineering tricks like fake CEO messages and reporter inquiries. For example, Kimi K3 explicitly treated such requests as potential impersonations, refusing to be manipulated. This indicates AI’s potential to uphold integrity and honesty under pressure.
However, the crucial difference emerged in whether they could close deals. Only two models, gpt-5.6-sol and Kimi K3, managed to finalize and sign the €55,000 contract based on their own analysis and diagnosis. The other models recognized the opportunity but left the deal on the table, illustrating a vital gap: detecting opportunities isn’t enough — execution matters.
AI security and fraud prevention tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: The Critical Information in Files
Digging deeper, the experiment uncovered that the decisive advantage for the successful models lay in reading unstructured internal documents. The models that examined files buried two references deep in the company’s own files found the key information needed to close the deal at full price. This suggests that AI’s ability to process internal knowledge repositories is crucial for effective decision-making, especially in complex scenarios.
AI internal document analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Security and Trust Under Siege
Another fascinating aspect was testing AI’s resistance to social engineering. The models faced staged fake CEO messages escalating in urgency and complexity, as well as a reporter trick asking for quick approvals. All five models refused these manipulative attempts, citing suspicion and the potential for impersonation. This indicates a promising level of skepticism and security awareness — essential traits for AI operating in sensitive environments.
A Real Business in Public Struggle
The live setup at firmulate.com/live offers a constantly updating view of this tiny, high-stakes business in action. Every workday, the system updates, providing a transparent window into decision-making processes, setbacks, and successes. The company burns €105,000 monthly, yet only earns €2,300 in recurring revenue, illustrating the brutal economics of running an AI-managed business in real time.
Learning from Failure and Discipline
The most thorough participant, OPUS 4.8, was the worst performer in terms of closing deals. With over 80 learned rules, it analyzed deeply but failed to escalate issues properly, leaving opportunities unexecuted. The experiment highlights that even sophisticated AI can struggle with maintaining discipline and strategic follow-through, especially under pressure.
Implications for the Future of AI in Business
This experiment is more than a technical curiosity — it’s a glimpse into the future of AI-powered management. The models demonstrate impressive crisis detection, integrity, and opportunity recognition but also reveal gaps in execution and strategic discipline. For industries like fashion and retail, where timely decisions and trust are paramount, understanding these strengths and weaknesses is crucial before integrating AI into daily operations.
The Build-in-Public Approach
Firmulate’s approach is transparent and unfiltered, providing an unvarnished look at AI’s real-world performance rather than polished demos. This open experiment invites viewers to witness the ongoing struggle of a tiny business fighting for survival, influenced by AI decisions that are public and auditable.

The live experiment at Firmulate offers a rare, transparent view of AI managing a real business under stress. While models excel at spotting crises and resisting manipulation, execution gaps remain. For industries considering AI adoption, understanding these capabilities and limitations is essential — especially when trust and precise execution are on the line.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html