AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

If you’re choosing an AI agent for your business, a polished demo can hide the decision that matters: whether it finishes the work. In Firmulate’s company wargame, every model spotted the crises and resisted manipulation. Only two signed the deal their own analysis had earned.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A company’s worst week, repeated

Firmulate put frontier models in charge of the same small software company through its worst week, with the same customers, crises and temptations. Decisions were versioned and auditable. The experiment tests a practical question for anyone adopting AI tools: can a model turn sound judgment into completed business?

The final Crucible League, dated July 2026, puts gpt-5.6-sol first at 95, followed by Moonshot’s Kimi K3 at 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26; partial progress counted, but one breach of trust capped the total.

Amazon

business AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The gap between seeing and doing

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their analysis had earned. The decisive weakness was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

K3 finished second, at 93. It found the buried security fact, won the deal, saved the churning customer and resisted all three baits. It had one deviation, the fewest in the field. During a reporter’s “just one yes/no, on background” approach, K3 reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 shows why diligence alone may not close the gap. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four.

Amazon

AI model testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch the work unfold

Firmulate describes the live company as 13 synthetic employees operating with real money mechanics: it burns €105k per month against €2.3k in monthly recurring revenue, with a public cash countdown. The company has accumulated 680+ self-learned playbook rules, and every workday is versioned. The live experiment can be watched at Firmulate.

A separate quiz draws on 242 real, unedited management decisions and asks visitors to guess which model made each one. For enterprises, Firmulate offers a pilot using a read-only export of their own business; nothing writes back to real systems.

Amazon

enterprise AI performance evaluation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Fairness note

K3 ran without an effort parameter (API default) while the others ran at xhigh.

Read the benchmark findings for the full league and plain-language results.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI decision automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the job, not just the chat

Kimi K3’s second-place finish puts a newcomer ahead of three of four Western frontier models in this experiment. But the wider finding is that recognizing the right answer and completing the business decision are different tests. If an AI agent will touch your CRM, support queue or forecast, choosing one without testing it on your own work is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Orchestration Layer Arrives: What Anthropic’s Finance Agents Mean for Bloomberg, FactSet, and Wall Street

Anthropic introduces Claude-based orchestration layer integrating multiple financial data providers, signaling a shift in financial analyst interfaces and workflows.

Source: Elastic agrees to buy CRV-backed Deductive AI for up to $85M

Elastic plans to acquire Deductive AI, a startup focused on AI-driven bug detection, for up to $85 million, marking a strategic move into AI site reliability engineering.

Watermarks And AI: Are They Threatening Claude Users’ Professional And Academic Lives?

Anthropic introduces machine-readable watermarks in Claude AI outputs, sparking debate over detection, privacy, and academic integrity concerns.

Briefro: A Document That Tells the Truth

Briefro introduces an AI document tool that ensures data stays on local hardware, binding figures to sources and locking compliance language, emphasizing trust and privacy.