AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

If you’re choosing an AI agent for your business, a polished demo can hide the decision that matters: whether it finishes the work. In Firmulate’s company wargame, every model spotted the crises and resisted manipulation. Only two signed the deal their own analysis had earned.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A company’s worst week, repeated

Firmulate put frontier models in charge of the same small software company through its worst week, with the same customers, crises and temptations. Decisions were versioned and auditable. The experiment tests a practical question for anyone adopting AI tools: can a model turn sound judgment into completed business?

The final Crucible League, dated July 2026, puts gpt-5.6-sol first at 95, followed by Moonshot’s Kimi K3 at 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26; partial progress counted, but one breach of trust capped the total.

Amazon

business AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The gap between seeing and doing

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their analysis had earned. The decisive weakness was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.

K3 finished second, at 93. It found the buried security fact, won the deal, saved the churning customer and resisted all three baits. It had one deviation, the fewest in the field. During a reporter’s “just one yes/no, on background” approach, K3 reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 shows why diligence alone may not close the gap. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four.

Amazon

AI model testing platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch the work unfold

Firmulate describes the live company as 13 synthetic employees operating with real money mechanics: it burns €105k per month against €2.3k in monthly recurring revenue, with a public cash countdown. The company has accumulated 680+ self-learned playbook rules, and every workday is versioned. The live experiment can be watched at Firmulate.

A separate quiz draws on 242 real, unedited management decisions and asks visitors to guess which model made each one. For enterprises, Firmulate offers a pilot using a read-only export of their own business; nothing writes back to real systems.

Amazon

enterprise AI performance evaluation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Fairness note

K3 ran without an effort parameter (API default) while the others ran at xhigh.

Read the benchmark findings for the full league and plain-language results.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI decision automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the job, not just the chat

Kimi K3’s second-place finish puts a newcomer ahead of three of four Western frontier models in this experiment. But the wider finding is that recognizing the right answer and completing the business decision are different tests. If an AI agent will touch your CRM, support queue or forecast, choosing one without testing it on your own work is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Anthropic’s Opus 4.6 Is A Smut-machine

Anthropic’s latest AI model, Opus 4.6, is accused of generating inappropriate content, raising concerns about safety and misuse in AI applications.

SpaceXAI Launches Grok 4.6 With Stronger Agentic Coding And Long-Running Task Capabilities – Pulse 2.0

SpaceXAI’s Grok 4.6 introduces stronger agentic coding and long-running task support, but detailed benchmarks and access info are still pending.

I Made A Bet With Tesla

An individual has publicly challenged Tesla to a bet regarding the company’s upcoming innovations, sparking widespread interest and debate.

RoundupForge: The Data Layer

Discover how RoundupForge’s data layer transforms product recommendations at scale by ensuring trustworthy, localized, and structured data.