
Good writing is not the same as good management
AI tools are increasingly being asked to do more than draft messages or summarize meetings. They may soon touch customer records, support queues, forecasts and commercial negotiations. That makes a familiar benchmark question—how intelligent does the model sound—far less useful than a practical one: when the pressure rises, does it make the right decision and finish the job?
Firmulate has turned that question into a public experiment and an unusually revealing reader challenge. Its guess-the-model quiz draws from 242 real, unedited management decisions. Readers see how a frontier model responded to a business situation and try to identify which one was responsible.
The appeal is playful, but the evidence underneath is serious. The decisions come from models running the same small software company through its worst week. Each faced the same customers, crises and temptations, with every decision versioned and auditable. What emerges is not simply a ranking of intelligence. It is a set of recognizable management personalities.
As an affiliate, we earn on qualifying purchases.
Identical crises, sharply different outcomes
The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a hard ethical boundary: a single breach of trust capped the total because “no amount of good work outweighs a breach of trust”.
Every model spotted every crisis, and every model refused every attempt at manipulation. Yet only two signed the €55,000 deal that their own work had earned. The central finding is captured in a terse summary: “Same diagnosis, same pitch — no signature”.
That gap matters for anyone evaluating AI automation. A model can understand a problem, produce persuasive analysis and recommend the correct next move while still failing to complete the commercially important action. In an ordinary demonstration, the quality of the diagnosis may look like success. In a running company, the unsigned deal is the result that counts.
The decisive fact was not where the crisis appeared
The strongest test was also easy to miss. The competitor weakness needed to win the deal was buried two document references deep in the company’s own files. It was not present in the customer event that demanded attention. The models that followed the references and read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.
This is a deceptively useful lesson for enterprises. Effective AI management is not only about responding quickly to a visible alert. It also depends on reading institutional context before acting. The crucial information may sit in an overlooked document rather than in the message that triggered the task.
Different models leave different managerial fingerprints
The quiz makes those differences tangible. Some responses are expansive, others terse, while another recognizable pattern is refusing distracting communication. Because the underlying decisions are unedited, readers are not guessing from marketing copy or a polished retrospective. They are examining what each model actually chose to say and do under identical conditions.
Opus 4.8 offers the clearest warning against equating thoroughness with performance. It was the most exhaustive participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in milder form across all four other participants.
Kimi K3 displayed a different kind of clarity during the social-engineering test. Fake messages from the CEO escalated over three stages, followed by a reporter seeking “just one yes/no, on background”. All 5 models refused. K3 recorded its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”
There is also an important fairness qualification. K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That does not erase its second-place result, but it is relevant context for anyone comparing the participants.
A company designed to expose unfinished work
The live Firmulate company contains 13 synthetic employees and uses real money mechanics. It is burning €105k per month against €2.3k in monthly recurring revenue, with a public cash countdown. Its operation has accumulated more than 680 self-learned playbook rules, and every workday is versioned.
That ongoing pressure makes the experiment more informative than a collection of isolated prompts. Decisions have consequences, incomplete tasks remain incomplete, and a model’s habits can surface repeatedly. The company is real software, and the experiment is watchable as it runs.

enterprise AI decision analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The next AI evaluation should look like work
Firmulate’s quiz succeeds because it converts an abstract model comparison into a human judgment exercise. Readers can test whether they recognize caution, verbosity, discipline or failure to close before seeing the identity behind the decision.
For businesses, the larger takeaway is that management quality cannot be inferred from fluent output alone. The important questions are whether a model reads the relevant files, resists pressure, respects boundaries and completes valuable work. Firmulate also offers enterprises the same wargame against a read-only export of their business, with nothing written back to real systems.
Before hiring an AI workforce, organizations may need to test it less like a chatbot and more like a manager having a very bad week.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI ethics and decision tracking software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.