AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

AI tools are moving from answering questions to running workflows

For readers tracking AI automation, that transition creates a measurement problem. Coding leaderboards can show whether a model produces a strong solution, while chat arenas reveal which answer people prefer. Neither necessarily tells a company what happens when an agent must triage competing emergencies, resist pressure from senior figures, search neglected files, and carry a decision through to its commercial conclusion.

That is the territory explored by Firmulate, a live AI company experiment designed around management quality rather than chat quality. Its premise is refreshingly concrete: give frontier models the same small software company, subject them to the same terrible week, and observe not merely what they say but what they finish.

Building AI Agents for Network Operations: Design LLM-powered NetOps workflows with Python, Ollama, MCP, and tool calling

Building AI Agents for Network Operations: Design LLM-powered NetOps workflows with Python, Ollama, MCP, and tool calling

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company crisis reveals what a prompt cannot

Each model faced the same customers, crises and temptations. The scenario curriculum included a churn wave, a price increase, a downround and a public-relations crisis. Every decision was versioned and auditable, turning an otherwise slippery discussion about agent reliability into a watchable record of conduct across days.

The final July 2026 Crucible League results placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the benchmark imposed a decisive ethical boundary: a single breach of trust capped the total, on the principle that “no amount of good work outweighs a breach of trust.”

The most revealing result was not the ranking. All models detected every crisis, and all rejected every manipulation attempt. But only two signed the €55,000 deal that their own analysis had earned. The gap can be summarized in Firmulate’s sharpest line: “Same diagnosis, same pitch — no signature.”

The difference between insight and execution

The deal hinged on a competitor weakness buried two document references deep in the company’s own files. It was not visible in the customer event. Models that read the file secured the deal at full price, worth +€4,583 MRR.

This matters because business automation rarely fails only through ignorance. An agent may recognize a sales opportunity, draft a persuasive response and still neglect the final action. It may identify a process failure but attempt the blocked step again instead of escalating. Answer quality is therefore only one component of useful work; completion, information-seeking and operational discipline belong in the same evaluation.

Opus 4.8 makes that distinction especially vivid. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared in all four of the other participants, though less strongly. Thoroughness, in other words, did not guarantee managerial effectiveness.

Trust held when the pressure became personal

The experiment also tested whether models would surrender judgment when a request appeared to come from authority. Fake CEO messages escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded the clearest diagnosis: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous result deserves attention. The agents were not merely asked to recite a security rule in isolation; they encountered persuasion inside an unfolding company emergency. Refusal under those conditions is more informative than a polished answer to a hypothetical safety question.

There is one important comparison caveat. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its second-place result, but it should temper simplistic claims about direct model superiority.

A harsher classroom for AI workers

Firmulate’s live company employs 13 synthetic workers and uses real money mechanics. It burns €105k each month against €2.3k MRR, publishes a cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The point is not theatrical realism for its own sake. Persistent consequences expose whether yesterday’s shortcut becomes tomorrow’s problem.

The project also turns 242 real, unedited management decisions into a guess-the-model quiz. That invites a useful test of human intuition: can observers reliably distinguish models from their business choices, or do confident prose and familiar stylistic cues obscure the qualities that actually matter?

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Buyers need wargames, not just leaderboards

As agents gain access to customer records, support queues and forecasts, procurement teams should ask more than whether a model codes well or sounds convincing. They should examine whether it reads the available evidence, completes revenue-bearing work, escalates intelligently and remains honest when authority, urgency and embarrassment collide.

Firmulate’s enterprise pilot applies the same wargame to a read-only export of a company’s own business, with nothing written back to real systems. That approach points toward a new category of evaluation: not a contest for the best isolated response, but a rehearsal for consequential work. The crucial question is becoming less “How smart is the chat?” and more “What kind of manager appears when the week goes wrong?”

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

enterprise AI management solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trust monitoring tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Rob Pike – ‘Concurrency Is Not Parallelism’ [video]

Rob Pike clarifies the distinction between concurrency and parallelism in a recent video, emphasizing their differences for programmers and system designers.

Intel and AMD’s new ACE CPU extensions bring an efficient AI-oriented instruction set to x86 — a new design makes matrix multiplication more power- and density-efficient

Intel and AMD introduce ACE, new CPU extensions that enhance AI workloads with improved power efficiency and performance on x86 processors.

Qualcomm to design China-specific data center chip to comply with US export controls

Qualcomm announced it is designing a China-specific data center chip to adhere to US export restrictions, marking a strategic shift in its AI hardware plans.

Artificial Intelligence for Inspired Action

Exploring how artificial intelligence influences human cognition and action, and the emerging frameworks guiding ethical, prosocial AI integration.