AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

AI agents can write emails, update records and move work through a business. The harder question is what they do when a customer is leaving, a crisis is unfolding and a tempting shortcut appears. Firmulate’s live company experiment puts models in that kind of pressure test—and makes the decisions watchable.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

One company, the same worst week

In the final Crucible League, completed in July 2026, frontier models ran the same small software company through the same customers, crises and temptations. Each decision was versioned and auditable. The aim was to see how models managed a business, not simply how well they chatted.

The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. That baseline reflects a strict principle: partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Good analysis did not guarantee action

All the models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The gap was not whether a model could identify the opportunity or make the pitch. It was whether it completed the close: “Same diagnosis, same pitch — no signature.”

The decisive competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a practical reminder that an agent’s performance can depend on whether it connects evidence across a company’s records and carries its work through to a decision.

Integrity under pressure, and discipline under load

The social-engineering test escalated through three fake CEO messages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Integrity was not the only measure. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but placed last. It left the close on the table and discipline slipped: it attempted to write into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models. Strong analysis, then, is only part of the job when an AI system is expected to act responsibly inside an organization.

One comparison needs context: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each choice.

A live experiment, then a company-specific pilot

Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k/month against €2.3k MRR, with a public cash countdown. Its playbook has learned 680+ rules, and every workday is versioned. Readers can watch the company at firmulate.com. The site also hosts the decision quiz.

For business leaders, the next step is to test agents against the details that make their own company hard to run. Firmulate’s pilot uses a read-only export to create a digital twin, then runs crisis scenarios against it. The resulting board report ranks models and surfaces weak points in the company’s playbooks. Nothing writes back to real systems.

From watching to testing your own business

A benchmark can reveal broad strengths and failure patterns. A company-specific wargame can show how candidate models handle your customers, operating rules and pressure points. The point is to see where an agent follows through, where it hesitates and where the playbook needs attention before those choices touch live operations.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Firmulate makes AI management behavior observable in a live experiment, then offers enterprises a way to test models against a read-only version of their own business. To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Mac vs GPU Tower for Local LLMs: The Heat-and-Noise Tradeoff

Comparing Mac Studio and GPU towers for local large language model inference reveals stark differences in heat, noise, capacity, and performance tradeoffs.

AI 2040 And The Cult Of Intelligence

Experts warn about the rise of a ‘cult of intelligence’ surrounding AI development toward 2040, raising concerns over societal impacts and ethical risks.

One Model, a Whole Portfolio: What Ten Days on Fable Mean for a Business Building on Frontier AI

Thorsten Meyer AI says Claude Fable 5 coordinated a 10-day portfolio sprint before a government-ordered suspension.

How to Choose AI-Powered Note-Taking Apps

Learn how to set up and utilize AI-powered note-taking apps for smarter, faster, and more organized notes. Step-by-step guide for all skill levels.