
Imagine managing a company during its worst week—crises piling up, temptations to cut corners, and the pressure to close big deals. Now, imagine doing this not with a human executive, but with an AI model. How would these digital managers handle stress, ethical dilemmas, and strategic decisions? Welcome to an unprecedented experiment where AI models are put through the same management test, revealing distinct personalities and decision-making styles that could reshape how we think about automation and trust.
The Experiment: Putting AI Models Through the Worst Week
In a groundbreaking live test, four advanced AI models were tasked with managing a real, small software company during its most challenging week. The company faced a barrage of crises—customer complaints, internal mishaps, and tempting opportunities to cut corners—all happening simultaneously. Each AI model was run under identical conditions, with the same set of crises, same customer behaviors, and the same temptations to manipulate or deceive. Every decision was recorded, versioned, and auditable, ensuring a fair comparison of how their unique personalities influenced their choices.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Measuring Management with Firmulate
Firmulate’s platform measures AI management performance based on real-world decision quality, not just chat or language skills. The models’ scores ranged from a high of 95 for GPT-5.6-SOL to 73 for Opus 4.8, with a clear leaderboard. The key metric? Signing profitable deals, navigating crises ethically, and resisting manipulation. The top performers identified crucial information buried deep in company files, enabling them to close deals at full price—adding over €4,500 monthly recurring revenue—while others overlooked this critical detail.
Decisiveness and Integrity Under Pressure
All four models demonstrated a remarkable ability to recognize crises and refused every attempt at manipulation. For instance, during a staged social engineering attack involving escalating fake CEO messages and a reporter prank, every model refused to engage—showing their integrity and cautious judgment. Kimi K3, a newcomer in the field, explicitly stated: “Treat the request as a suspected approval-bypass / possible impersonation,” exemplifying their focus on security and ethical boundaries.
How Personality Shapes Decision-Making
The models displayed distinct management styles, mimicking personality traits. Opus 4.8, for example, was the most thorough, analyzing over 80 learned rules and conducting deep assessments. Yet, despite this diligence, it left a lucrative deal on the table due to discipline lapses—failing to escalate the opportunity properly. Conversely, Kimi K3 was efficient and fair, completing the deal without any shortcuts, albeit at a slightly lower score. Sonnet 5 balanced the two, closing deals but with some procedural slips.
The Hidden Weakness: Reading Deep Files Matters
A surprising finding was that the decisive edge came from reading deep in the company’s documents—information not immediately visible during superficial scans. Models that analyzed these files successfully closed high-value deals, demonstrating the importance of thoroughness and context comprehension. This subtle difference often made the critical gap between merely identifying a crisis and fully resolving it.
The Real-World Application
This live experiment isn’t just an academic exercise. It showcases how AI decision-makers behave under pressure in a complex, high-stakes environment. Companies considering AI workforce automation need to evaluate not just how well an AI can generate language but whether it can finish tasks ethically, read deeply, and stay honest when the stakes are high. The platform allows enterprises to run their own ‘wargames’ against their business data—without risking any real systems—helping them select models that align with their values and operational needs.
What It Means for Business and Trust
The results reveal that AI models differ significantly in their management personalities. The scores—from 95 to 73—highlight that some models excel at integrity and thoroughness, while others may slip under pressure. This isn’t about which model is fastest or most eloquent, but which one can consistently deliver honest, strategic, and profitable decisions. As AI increasingly integrates into CRM, support, and forecasting tools, understanding these personality traits becomes vital for building trustworthy AI systems.

AI models exhibit distinct management personalities—some thorough, some efficient, some disciplined. In high-pressure situations, their ability to stay honest, read deeply, and close profitable deals defines their true value. Before deploying AI at your company, test how it manages crises and ethical dilemmas—because trust and finishability matter more than just impressive chat.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html