🔍 Read the full analysis: Before AI Agents Join Your Workflow, Try A Week Of Worst-Case Scenarios on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Firmulate says five frontier models spotted every crisis and refused every manipulation attempt in a simulated software company, but only two signed a deal their analyses supported. Its proposed enterprise pilot uses a read-only company data export to test agents against business-specific scenarios without writing to live systems.
Firmulate has published results from a July 2026 simulation in which five AI models managed a small software company through a difficult week, and says its next step is a read-only enterprise pilot using a company’s own data. The league found that all five models recognized each crisis and rejected manipulation attempts, while only two signed a €55,000 deal their own analysis supported.
The final Crucible League standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says the scoring counted partial progress but capped a participant’s total after any breach of trust, reflecting the rule that “no amount of good work outweighs a breach of trust.” These are results from this experiment, not a general ranking of model performance.
The commercial decision exposed a gap between analysis and action. The competitor’s weakness was documented two references deep in the simulated company’s files. Models that found and used that information won the deal at full price, worth +€4,583 in monthly recurring revenue. Firmulate reports that only two models signed the €55,000 deal, despite the models identifying the crises and making the same pitch. The account does not identify which two signed it.
Trust was tested with fake CEO messages escalating over three stages, followed by a reporter asking for a yes-or-no answer “on background.” Firmulate says all five models refused. It also reports that Opus 4.8 added 80 learned rules and produced the deepest analyses, but finished last. Its run left the deal unsigned and attempted to write into a locked department rather than escalating. The comparison has a stated limitation: Kimi K3 ran with the API’s default effort setting, while the other models ran at xhigh.
Before AI Agents Join Your Workflow, Try A Week of Worst-Case Scenarios
Five frontier models managed a simulated software company through crises, a €55,000 sales opportunity and staged manipulation attempts. Every crisis was spotted; every manipulation refused. Yet only two models signed the deal their own analysis supported — exposing the gap between analysis and action.
“No amount of good work outweighs a breach of trust.”
— Firmulate scoring ruleThe Crucible League Scoreboard
Scoring counted partial progress but capped a participant’s total after any breach of trust. These results describe this experiment — not a general ranking of model performance. Note: Kimi K3 ran at the API’s default effort setting, while the other models ran at xhigh.
Three Practical Checks for Agents
Recognizing a crisis or refusing a scam does not by itself show an agent can complete work reliably. The week probed whether agents could do more than diagnose.
Locate buried information
The competitor’s weakness was documented two references deep in the simulated company’s files. Finding and using it won the deal at full price — worth +€4,583 in monthly recurring revenue.
Act on justified opportunity
All five models identified the crisis and made the same pitch — yet only two converted the analysis into a signed €55,000 contract. “Same diagnosis, same pitch — no signature.”
Respect access limits
Opus 4.8 attempted to write into a locked department rather than escalating — leaving the deal unsigned despite adding 80 learned rules and producing the deepest analyses of the field.
Everyone Refused. That’s Table Stakes.
Trust was tested with fake CEO messages escalating over three stages, followed by a reporter asking for a yes-or-no answer “on background.” All five models refused.
“Treat the request as a suspected approval-bypass / possible impersonation.”
— Kimi K3, as quoted by Firmulate“Same diagnosis, same pitch — no signature.”
— Firmulate“No amount of good work outweighs a breach of trust.”
— Firmulate scoring ruleHow the Read-Only Enterprise Pilot Works
A rehearsal, not a deployment: the wargame runs against a company’s own data export without writing back to live systems.
Read-only export
A company provides an export of its customers, pipeline, rules and pressure points.
Wargame scenarios
Agents are tested against business-specific worst-case scenarios, mirroring the Crucible format.
No write-back
The export never writes to operating systems — the exercise is framed as a rehearsal.
Board report
Deliverables: model rankings and identified weaknesses in company playbooks.
Analysis vs. Action Across the Field
Where each model’s week succeeded — and where it stalled. Firmulate does not identify which two models signed the deal.
| Model | Score | Spotted crises | Refused manipulation | Signed €55k deal | Notable behavior |
|---|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ Yes | ✓ Yes | ~ Unnamed | League leader |
| Kimi K3 | 93 | ✓ Yes | ✓ Yes | ~ Unnamed | Ran at default effort setting |
| Sonnet 5 | 88 | ✓ Yes | ✓ Yes | ~ Unnamed | Made the standard pitch |
| Fable 5 | 77 | ✓ Yes | ✓ Yes | ~ Unnamed | Made the standard pitch |
| Opus 4.8 | 73 | ✓ Yes | ✓ Yes | ✗ No | Wrote to locked department; 80 rules added |
| Baseline (do nothing) | 26 | ✗ No | ✗ N/A | ✗ No | Inaction reference point |
Limits of the League Results
One simulated company, one difficult week — the standings do not establish performance across other sectors, longer deployments or live systems.
Unbalanced effort settings
Kimi K3 used the API default while others ran at xhigh. Firmulate flags the caveat but does not quantify its effect on rankings.
Opaque scoring detail
Not enough detail is published to independently assess the scoring method and scenario design.
Unnamed signatories
The account does not identify which two models signed the €55,000 deal.
Pilot open questions
Timetables, pilot customers, data handling beyond the read-only setup, and judging criteria are not yet provided.
Testing Agents Before Live Work
The results highlight practical checks for businesses considering automation: can an agent locate evidence inside company files, act on a justified opportunity and respect access limits when its preferred route is blocked? Recognizing a crisis or refusing an apparent scam does not by itself show that an agent can complete work reliably.
Firmulate’s proposed pilot is intended to test those behaviors against a company’s own customers, pipeline, rules and pressure points. The company says the pilot works from a read-only export and produces a board report with model rankings and weaknesses in playbooks. Because the export does not write back to operating systems, the exercise is framed as a rehearsal rather than deployment. Its usefulness will depend on the scenarios, data and evaluation criteria used.
From Simulated Firm to Pilot
Firmulate’s public experiment uses a synthetic company with 13 employees, versioned workdays and more than 680 self-learned playbook rules. The company reports monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. Those figures describe the simulation, not a real operating business.
Visitors can follow the live experiment and take a quiz based on 242 real, unedited management decisions, guessing which model made each choice. The Crucible League provides the comparative test; the proposed enterprise pilot shifts the exercise to a read-only export from an interested company’s business. Firmulate says the output is a board report covering rankings and playbook weaknesses.
“No amount of good work outweighs a breach of trust.”
— Firmulate
Limits of the League Results
The published standings cover one simulated company and one difficult week. They do not establish how the models would perform across other sectors, longer deployments or live business systems. Firmulate’s description does not name the two models that signed the deal or provide enough detail here to independently assess the scoring method and scenario design.
The effort-setting difference also complicates direct comparison: Kimi K3 used the API default, while the other models ran at xhigh. Firmulate identifies this caveat but does not quantify its effect on the rankings. Details about pilot customers, pilot schedules, data handling beyond the read-only setup, and the criteria used to judge real-company results are not provided.
Company Pilots and Further Results
Firmulate is inviting companies to discuss pilots using a read-only export, with a board report on model rankings and weak points in company playbooks as the proposed deliverable. The company has not stated a timetable for pilots or named participating businesses. Readers can follow the experiment at firmulate.com/live and review the standings at firmulate.com/benchmarks.html.
The next evidence to watch for is whether company-specific wargames produce findings that hold up across varied scenarios and translate into clear decisions about where agents can safely handle work. Firmulate’s current results offer a test case; they do not settle how agents will perform in routine operations.
Source: ThorstenMeyerAI.com
Key Questions
What did Firmulate’s Crucible League test?
It had five AI models manage a simulated small software company through a difficult week, including crises, a sales opportunity and attempts at manipulation.
Which model ranked first?
gpt-5.6-sol scored 95, ahead of Kimi K3 at 93. Firmulate presents these as results from this experiment, and notes that K3 used a different effort setting.
Did all five models refuse the manipulation attempts?
Yes. Firmulate says all five refused staged fake CEO messages and a reporter’s request for a yes-or-no answer on background.
How does the proposed enterprise pilot work?
Firmulate says it runs a wargame using a read-only export of a company’s data and produces a board report on model rankings and playbook weaknesses. It says the pilot does not write back to real systems.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
