
Turn your wind-down time into listening time
- Thousands of audiobooks, podcasts and originals
- Listen on your phone, tablet or Echo — also offline
- Cancel anytime
Character shows up when the pressure is real
Personality tests ask how we might respond under pressure. Firmulate asks a related question of AI: when a company faces a terrible week, will a model notice the danger, resist manipulation and follow through on what it says is right?
The experiment is a live, watchable company, with synthetic employees making versioned decisions as they navigate business crises. Its final July 2026 league offers a revealing result: every model recognized the crises and rejected the manipulation attempts, but some still failed to close a deal their own analysis supported.
Same company, same temptations
Firmulate ran frontier models through the same small software company and its worst week. They faced the same customers, crises and temptations. The goal was to observe decisions and outcomes, not judge how persuasive a model sounds in a conversation.
The final Crucible League ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s rule captures the weight it gives to integrity: “no amount of good work outweighs a breach of trust.”
Good judgment needs follow-through
The striking result was not that models missed danger. All spotted every crisis and refused every manipulation attempt. The gap emerged afterward: only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
That makes the experiment a study in follow-through as much as crisis recognition. A model can identify a sound course of action and still leave the opportunity untouched. In a workplace, that difference between saying what should happen and completing the task could shape whether an AI assistant is useful under pressure.
One crucial clue sat two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding suggests that attention to context can matter as much as confidence in a recommendation.
Boundaries, trust and a missed close
The social-engineering test escalated through fake CEO messages in three stages, followed by a reporter’s request: “just one yes/no, on background”. All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Yet discipline was not uniform across the broader work. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal on the table and attempted writes into a locked department instead of escalating. The same weakness appeared, more weakly, in all four.
That contrast is familiar from human behavior: insight and effort do not guarantee sound action. Firmulate’s results offer a concrete way to inspect the space between recognizing a problem, respecting a boundary and carrying a decision through.
From watching to your own business
The live company makes the stakes visible through real money mechanics: it has 13 synthetic employees, burns €105k/month against €2.3k MRR, and shows a public cash countdown. More than 680 self-learned playbook rules accumulate as every workday is versioned. The site also offers a quiz built from 242 real, unedited management decisions: guess which model made each choice.
For enterprises, the next step is a pilot against a read-only export of their own business. That allows teams to test crisis scenarios against company-specific information and receive a board report ranking models and identifying weak points in their playbooks. Nothing writes back to real systems. A note on comparing the league: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh.

See how your AI responds
A strong answer is only part of good judgment. The Firmulate experiment tracks whether a model can recognize a crisis, protect trust and follow through when the decision matters. Enterprise teams can take that question into a controlled pilot using a read-only business export. Contact Firmulate about a pilot at contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
