AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

An experiment by Firmulate demonstrates that AI models can diagnose crises and formulate responses but often fail to finalize work or close deals under operational pressures. This exposes a gap between AI understanding and execution in business contexts.

Firmulate’s recent live experiment revealed that while AI models can accurately diagnose crises and generate appropriate responses, they often fail to complete the necessary steps to finalize trust-based work, such as closing deals or executing decisions, under real operational pressures. This exposes a gap between AI understanding and actual execution in business environments, with significant implications for AI’s management challenges in enterprise management.

The experiment involved a simulated company with 13 synthetic employees and real financial metrics, where AI models were tasked with handling crises, identifying critical information buried deep in documents, and making operational decisions. For more context, see the original analysis on AI’s management gap. Despite all models recognizing crises and resisting manipulation attempts, only two successfully signed a €55,000 deal, illustrating that correct diagnosis does not guarantee successful completion of work. This highlights the importance of operational discipline, as detailed in the original analysis.

Notably, the models that achieved the highest scores in the benchmark, such as gpt-5.6-sol, demonstrated strong understanding but still showed a clear divide when transitioning from analysis to action. For example, one model identified a hidden document reference that led to a deal but failed to follow through with the final signature, highlighting the distinction between knowledge and execution discipline.

The experiment also tested models against social-engineering attempts, such as fake CEO messages, which all models recognized and refused, indicating safety awareness. However, the most thorough analysis did not necessarily translate into successful business closure, as demonstrated by Opus 4.8, which, despite deep analysis, failed to finalize a deal due to lapses in operational discipline.

At a glance
reportWhen: ongoing; results published in July 2026
The developmentFirmulate’s live company experiment tested AI models’ ability to turn correct analysis into completed, trustworthy work, revealing management and execution failures despite accurate diagnostics.

Implications for AI Adoption in Business Decision-Making

This experiment underscores a critical challenge for enterprises adopting AI: the importance of execution discipline. While models can diagnose and reason effectively, their ability to complete high-stakes tasks reliably remains limited. The distinction between understanding and doing is vital, as a correct analysis that is not followed through can result in costly failures, even when the AI’s reasoning is sound. For decision-makers, this highlights the need to evaluate not just AI accuracy but also its capacity to deliver finished, trustworthy work in operational settings.

Building AI Agents for Network Operations: Design LLM-powered NetOps workflows with Python, Ollama, MCP, and tool calling

Building AI Agents for Network Operations: Design LLM-powered NetOps workflows with Python, Ollama, MCP, and tool calling

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Gap Between AI Diagnostics and Operational Completion

Recent developments in AI have focused heavily on diagnostic and reasoning capabilities, with many models demonstrating proficiency in summarization, analysis, and safety recognition. However, the practical application of these models in real-world business processes—such as closing deals, executing decisions, or managing crises—remains a challenge. The Firmulate experiment builds on prior work by testing AI in a controlled but realistic environment, exposing the disconnect between understanding and acting, which has been a persistent concern in enterprise AI deployment.

Previous benchmarks and research have highlighted AI’s strengths in reasoning but have often overlooked the importance of operational discipline—an area now shown to be a critical factor in determining real-world success or failure.

“The models understood the situation and formulated the right response, but completing the work was a separate challenge.”

— an anonymous researcher

Amazon

AI business decision execution software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About AI Operational Reliability

It remains unclear whether these findings are specific to the experimental setup or indicative of broader limitations in current enterprise AI systems. The extent to which models can be trained or designed to improve their execution discipline, especially in high-pressure situations, is still being explored. Additionally, the impact of different operational frameworks or human oversight on AI performance in completing work is not yet fully understood.

Amazon

AI deal closing automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Evaluating and Improving AI Execution

Organizations are advised to conduct similar live tests or simulations tailored to their specific workflows to assess AI’s ability to transition from diagnosis to action. Further research will likely focus on integrating operational discipline into AI training, developing better oversight mechanisms, and establishing benchmarks that measure not only understanding but also successful completion of critical tasks. Monitoring how AI performs under real-world pressures will be essential for safe and reliable deployment.

Amazon

enterprise AI management solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why do AI models often fail to complete work despite understanding it?

While models can diagnose and reason effectively, completing work requires operational discipline and the ability to act within organizational constraints, which may not be fully developed in current AI systems.

What does this experiment reveal about AI safety and trust?

It shows that safety awareness alone does not guarantee successful execution. Trust depends on consistent performance in completing tasks, not just recognizing risks or generating correct responses.

How can organizations improve AI’s ability to finalize work?

They should evaluate AI performance through live simulations, focus on operational discipline, and develop oversight mechanisms to ensure AI actions align with organizational goals and constraints.

Is this issue specific to certain AI models or general across all systems?

The findings suggest a broader challenge in current AI architectures, but further testing across different models and contexts is needed to determine the generality of these results.

What are the risks of relying on AI that can diagnose but not execute?

Relying solely on diagnostic capabilities can lead to situations where issues are identified but not resolved, resulting in missed opportunities, failed deals, or unmitigated crises, especially under pressure.

Source: ThorstenMeyerAI.com

You May Also Like

Expanding OpenAI’s Presence In Brazil

OpenAI announced its official commercial launch in Brazil on August 27, establishing a São Paulo-based team to work with local businesses, researchers, and institutions.

Openai’s Global Hunt for GPU Power Sends Sam Altman Traveling Nonstop

I’m fascinated by OpenAI’s relentless quest for GPU dominance, as Sam Altman’s nonstop travels reveal a race to reshape AI infrastructure worldwide.

NotebookLM Is Now Gemini Notebook

Google has rebranded its NotebookLM AI tool as Gemini Notebook, signaling a shift in branding and development focus. Details remain limited.

Lithuanian startup launches open-source network to detect Shahed-type drones

A Lithuanian startup has launched an open-source system using volunteers’ smartphones to detect Shahed-type drones, aiming to enhance regional security.