AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

What if an AI business tool had to live with its decisions?

For readers following AI tools and automation, polished demos are becoming less persuasive. The harder question is what happens after an agent enters a messy workplace: Does it notice danger, resist manipulation, examine the available evidence and complete commercially useful work?

Firmulate is making that question unusually concrete. Its live software company has 13 synthetic employees, burns €105k each month against €2.3k in monthly recurring revenue and displays a public cash countdown. Its workforce has accumulated more than 680 self-learned playbook rules, while every workday is versioned. The result is a business experiment that unfolds as an observable corporate survival story, with new decisions and consequences appearing through the ordinary rhythm of work.

Anyone can watch the company live. That openness turns an abstract debate about autonomous AI into something closer to ongoing business reporting: the company has customers, financial pressure and a shrinking margin for mistakes.

AI Co-Thinking: A Framework for Working with AI

AI Co-Thinking: A Framework for Working with AI

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company under pressure becomes a benchmark

The live operation supplies the larger setting, but Firmulate’s Crucible League isolates the question of management quality. Each frontier model was asked to run the same small software company through its worst week. The customers, crises and temptations were held constant, and every decision was versioned and auditable.

The final July 2026 standings put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. There was also a hard constraint on misconduct: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The broad result was reassuring. Every model identified every crisis, and all of them rejected every attempt at manipulation. Yet recognition was not the same as execution. Only two models signed the €55,000 deal their own work had justified. The experiment’s sharpest summary is: “Same diagnosis, same pitch — no signature.”

The decisive information was not in the obvious place

The difference hinged on research rather than eloquence. A competitor’s crucial weakness was buried two document references deep in the company’s own files, rather than presented in the customer event. Models that followed the trail found the evidence and won the deal at full price, adding €4,583 in monthly recurring revenue.

That finding should resonate with companies evaluating agents for sales, support or operational work. A model can understand the immediate request and still miss the fact that determines the outcome. Good language is visible at once; disciplined investigation is harder to demonstrate until the agent is operating inside a realistic business situation.

Pressure also tested judgment and trust

The models faced fake messages from the chief executive that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

This matters because automation changes the scale at which social engineering can cause damage. An AI worker may encounter requests that sound urgent, authoritative or harmless while actually attempting to evade established approval. In Firmulate’s worst week, the entire field maintained the boundary.

K3’s strong performance carries an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference does not erase the result, but it belongs beside the ranking when readers compare participants.

Thoroughness did not guarantee completion

Opus 4.8 offers the most instructive counterexample. It was the most thorough participant, producing the deepest analyses and learning 80 additional rules, yet it finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating the issue.

A weaker version of that same problem appeared in all four other models. The lesson is not that analysis lacks value. It is that enterprise automation must join analysis to follow-through, procedural judgment and escalation. Firmulate also publishes the synthetic workforce’s own words, giving observers another view of how decisions are framed under pressure.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Building AI Solutions with Azure AI: Enterprise Applications and Intelligent Automation

Building AI Solutions with Azure AI: Enterprise Applications and Intelligent Automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The real product is an observable management record

Firmulate’s live company pushes build-in-public beyond launch updates and revenue charts. It exposes a running organization’s workdays, learned rules, financial strain and decision record while the outcome remains unsettled. The public cash countdown makes failure a present business possibility rather than a retrospective case study.

For buyers of AI tools, the experiment suggests a more demanding checklist:

  • Does the agent inspect company knowledge deeply enough to find buried commercial facts?
  • Can it distinguish urgent authority from an attempt to bypass approval?
  • Will it escalate correctly when a boundary blocks its intended action?
  • Does it finish valuable work after producing a convincing analysis?

Firmulate’s evidence shows why these questions belong together. The models could spot crises and resist manipulation, yet some still failed to complete the deal. In autonomous business software, trustworthy restraint and commercial follow-through are separate capabilities—and a company needs both.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions

The AI-Driven Leader: Harnessing AI to Make Faster, Smarter Decisions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

AI Automation for Local Businesses: No-Code Workflows to Get More Leads, Follow-Up Faster, and Increase Sales

AI Automation for Local Businesses: No-Code Workflows to Get More Leads, Follow-Up Faster, and Increase Sales

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

SenseTime Group, Inc. Class B Revenue Breakdown – HKEX:20 – TradingView

TradingView lists a revenue breakdown for SenseTime Group’s Class B shares under HKEX:20, but no financial figures or details are confirmed yet.

AI output review queue for customer support macros

Support teams are testing a new AI macro review queue to ensure policy compliance and tone consistency before publication.

Breakout Tools Launches MyBreakoutOS Mobile App With AI Agents And Knowledge Base – Carroll County Mirror-Democrat

Breakout Tools introduces MyBreakoutOS, a mobile app featuring AI agents and a knowledge base, expanding its platform for productivity and security.

How Much Does Sovereign AI Really Cost? Forge Or Self-Host?

Evaluating the true costs of sovereign AI: is building in-house cheaper than buying managed solutions? Key insights from recent industry analysis.