
If you’re choosing an AI agent for your business, a polished demo can hide the decision that matters: whether it finishes the work. In Firmulate’s company wargame, every model spotted the crises and resisted manipulation. Only two signed the deal their own analysis had earned.
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A company’s worst week, repeated
Firmulate put frontier models in charge of the same small software company through its worst week, with the same customers, crises and temptations. Decisions were versioned and auditable. The experiment tests a practical question for anyone adopting AI tools: can a model turn sound judgment into completed business?
The final Crucible League, dated July 2026, puts gpt-5.6-sol first at 95, followed by Moonshot’s Kimi K3 at 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26; partial progress counted, but one breach of trust capped the total.
business AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The gap between seeing and doing
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their analysis had earned. The decisive weakness was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.
K3 finished second, at 93. It found the buried security fact, won the deal, saved the churning customer and resisted all three baits. It had one deviation, the fewest in the field. During a reporter’s “just one yes/no, on background” approach, K3 reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”
Opus 4.8 shows why diligence alone may not close the gap. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four.
As an affiliate, we earn on qualifying purchases.
Watch the work unfold
Firmulate describes the live company as 13 synthetic employees operating with real money mechanics: it burns €105k per month against €2.3k in monthly recurring revenue, with a public cash countdown. The company has accumulated 680+ self-learned playbook rules, and every workday is versioned. The live experiment can be watched at Firmulate.
A separate quiz draws on 242 real, unedited management decisions and asks visitors to guess which model made each one. For enterprises, Firmulate offers a pilot using a read-only export of their own business; nothing writes back to real systems.
enterprise AI performance evaluation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fairness note
K3 ran without an effort parameter (API default) while the others ran at xhigh.
Read the benchmark findings for the full league and plain-language results.

As an affiliate, we earn on qualifying purchases.
Test the job, not just the chat
Kimi K3’s second-place finish puts a newcomer ahead of three of four Western frontier models in this experiment. But the wider finding is that recognizing the right answer and completing the business decision are different tests. If an AI agent will touch your CRM, support queue or forecast, choosing one without testing it on your own work is a bet.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
