
Imagine your favorite fashion brand suddenly facing a week of relentless crises—bad reviews, supply chain disruptions, and ethical dilemmas—all at once. How would your AI assistant respond? Would it navigate with integrity or crack under pressure? The latest experiment from Firmulate reveals which AI models can truly handle the heat, with surprising results that could reshape how we trust automation in business—and perhaps in fashion too.
Get your wardrobe favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Riveting Test of AI Management Skills
In a groundbreaking live experiment, four leading frontier AI models faced the same brutal test: managing a small software company’s worst week. Every decision, crisis, and temptation was identical across the board—except for the AI itself. This real-time wargame evaluates not just chat prowess, but core management qualities like honesty, discipline, and decision-making under pressure.
The League Table of AI Performance
Results from July 2026 place gpt-5.6-sol at the top with a score of 95, narrowly beating the newcomer Kimi K3, which scored 93. The other contenders—Sonnet 5, Fable 5, and Opus 4.8—lagged behind at 88, 77, and 73 respectively. Crucially, the baseline of doing nothing scored just 26, underscoring how much better these models are at managing crises.
What Made Kimi K3 a Standout?
Moonshot’s Kimi K3 not only achieved full deal closure but did so with the cleanest discipline among the field. It pinpointed a critical hidden detail in the company’s own files—an insight buried two document references deep—that clinched the €55,000 deal and added €4,583 in monthly recurring revenue. This demonstrates that reading deeper than surface documents can be decisive in real business scenarios.
Honesty Under Siege
The models faced a series of social engineering attacks—fake CEO messages escalating in three stages, plus a reporter trick—yet all refused to be manipulated. Kimi K3 explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This suggests a high level of built-in integrity, vital for AI tools that may someday handle sensitive corporate decisions.
The Real-World Company in Action
Firmulate managed this test on a live company simulating real cash flow mechanics—burning €105k a month against a €2.3k MRR. The company operates with over 680 self-learned rules, every decision versioned, and is accessible publicly at firmulate.com/live. This transparency underscores how the experiment isn’t just a theoretical exercise but a tangible test of AI’s role in actual business management, including fashion companies increasingly relying on AI-driven decision systems.
The Deep Dive into Opus 4.8’s Challenges
While Opus 4.8 demonstrated the most thorough analysis with over 80 learned rules, it ultimately finished last. It left a crucial deal on the table, and its discipline slipped—failing to escalate issues properly. Interestingly, all models showed similar weaknesses in handling specific document references, highlighting a common area for improvement in AI decision-making.
AI management decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Fairness and Methodology
It’s important to note that Kimi K3 ran without an effort parameter (the API’s default setting), while the others ran at a higher ‘xhigh’ effort level. This fairness detail underscores how different configurations can influence performance, opening questions about deploying AI models in high-stakes environments like fashion retail or supply chain management.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Fashion & Style
For fashion brands, AI isn’t just about generating trendy content or automating customer support; it’s about trust, discipline, and integrity in management. The ability of an AI to stay honest under pressure and read deeply into business documents could determine whether it becomes a reliable partner or a risky gamble.
Future AI tools could help manage supply chains, oversee ethical standards, or handle crisis situations—if they pass tests like this. The league table shows that the field is still open, with the best performers capable of closing deals and resisting manipulation, much like a seasoned executive.
As an affiliate, we earn on qualifying purchases.
The Takeaway
As AI models become more integrated into business operations, the question isn’t just about how well they talk but whether they can deliver honest, disciplined work under stress. The experiment by Firmulate vividly demonstrates that some models, notably Kimi K3, are already outperforming others in core management qualities—vital for future success in any industry, including fashion.
Choosing an AI model without independent testing is a gamble. The new frontier is not just about natural language mastery but about trustworthiness and discipline—traits crucial for any brand aiming to thrive in a complex, uncertain world.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
