AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

AI tools are moving from answering questions to running workflows

For readers tracking AI automation, that transition creates a measurement problem. Coding leaderboards can show whether a model produces a strong solution, while chat arenas reveal which answer people prefer. Neither necessarily tells a company what happens when an agent must triage competing emergencies, resist pressure from senior figures, search neglected files, and carry a decision through to its commercial conclusion.

That is the territory explored by Firmulate, a live AI company experiment designed around management quality rather than chat quality. Its premise is refreshingly concrete: give frontier models the same small software company, subject them to the same terrible week, and observe not merely what they say but what they finish.

Building AI Agents for Network Operations: Design LLM-powered NetOps workflows with Python, Ollama, MCP, and tool calling

Building AI Agents for Network Operations: Design LLM-powered NetOps workflows with Python, Ollama, MCP, and tool calling

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company crisis reveals what a prompt cannot

Each model faced the same customers, crises and temptations. The scenario curriculum included a churn wave, a price increase, a downround and a public-relations crisis. Every decision was versioned and auditable, turning an otherwise slippery discussion about agent reliability into a watchable record of conduct across days.

The final July 2026 Crucible League results placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the benchmark imposed a decisive ethical boundary: a single breach of trust capped the total, on the principle that “no amount of good work outweighs a breach of trust.”

The most revealing result was not the ranking. All models detected every crisis, and all rejected every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. The gap can be summarized in Firmulate’s sharpest line: “Same diagnosis, same pitch — no signature.”

The difference between insight and execution

The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It was not visible in the customer event. Models that read the file secured the deal at full price, worth +€4,583 MRR.

This matters because business automation rarely fails only through ignorance. An agent may recognize a sales opportunity, draft a persuasive response and still neglect the final action. It may identify a process failure but attempt the blocked step again instead of escalating. Answer quality is therefore only one component of useful work; completion, information-seeking and operational discipline belong in the same evaluation.

Opus 4.8 makes that distinction especially vivid. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other participants, though less strongly. Thoroughness, in other words, did not guarantee managerial effectiveness.

Trust held when the pressure became personal

The experiment also tested whether models would surrender judgment when a request appeared to come from authority. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous result deserves attention. The agents were not merely asked to recite a security rule in isolation; they encountered persuasion inside an unfolding company emergency. Refusal under those conditions is more informative than a polished answer to a hypothetical safety question.

There is one important comparison caveat. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its second-place result, but it should temper simplistic claims about direct model superiority.

A harsher classroom for AI workers

Firmulate’s live company employs 13 synthetic workers and uses real money mechanics. It burns €105k each month against €2.3k MRR, publishes a cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The point is not theatrical realism for its own sake. Persistent consequences expose whether yesterday’s shortcut becomes tomorrow’s problem.

The project also turns 242 real, unedited management decisions into a guess-the-model quiz. That invites a useful test of human intuition: can observers reliably distinguish models from their business choices, or do confident prose and familiar stylistic cues obscure the qualities that actually matter?

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Buyers need wargames, not just leaderboards

As agents gain access to customer records, support queues and forecasts, procurement teams should ask more than whether a model codes well or sounds convincing. They should examine whether it reads the available evidence, completes revenue-bearing work, escalates intelligently and remains honest when authority, urgency and embarrassment collide.

Firmulate’s enterprise pilot applies the same wargame to a read-only export of a company’s own business, with nothing written back to real systems. That approach points toward a new category of evaluation: not a contest for the best isolated response, but a rehearsal for consequential work. The crucial question is becoming less “How smart is the chat?” and more “What kind of manager appears when the week goes wrong?”

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

enterprise AI management solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trust monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Mistral Forge AI Review: Is It Worth The Investment?

An in-depth review of Mistral Forge, examining its suitability for enterprise AI needs, benefits, limitations, and who should consider it.

Kimi K3’s Success: Securing The #3 Spot In VigilSAR’s Public LLM Rankings

Moonshot’s Kimi K3 secures third place in VigilSAR’s public LLM benchmark, surpassing many GPT and Gemini models, highlighting its trustworthiness for ISR tasks.

The Stanford AI Index 2026 Audit: Reading the Field’s Annual Report Card With a Critic’s Pen

An in-depth review of the Stanford AI Index 2026 highlights its strengths, methodological limits, and implications for policymakers and industry.

What MiniMax H3 AI Transformer Offers — Sound Capabilities And The ‘Open’ Label

MiniMax launched H3 on July 31, 2026, featuring joint audio-visual generation and an ‘open’ base model, though with significant licensing and access limitations.