🔍 Read the full analysis: The Agent Said It Was Done. The Database Disagreed. on ThorstenMeyerAI.com
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Microsoft and Hugging Face have made ThinkingBox available through Hugging Face. The benchmark runs AI agents through 507 business workflows and checks whether backend records and side effects match each task’s requirements; its authors report that many attempts fail those checks despite making state-changing tool calls.
Microsoft and Hugging Face have made ThinkingBox available through Hugging Face, introducing a benchmark that tests whether AI agents leave business systems in the required state after completing tasks. Its authors say the benchmark covers 507 business workflows, each run up to 20 times, and checks records and side effects rather than relying only on valid tool calls or convincing responses.
ThinkingBox runs agents in isolated sessions using MCP tools, then checks the backend state at the end of each run. The benchmark spans retail, auto insurance, travel, neobanking and consulting. Repeating tasks is intended to measure consistency as well as whether an agent can succeed on an individual attempt.
In a common-set analysis, the authors report 121,680 valid trials across 12 models, of which 79,853 failed executable checks. Among those failed attempts, 67.24% ended without a final tool error despite the agent having invoked a state-changing tool. The authors found wrong field values in 77.61% of failures, unintended extra effects in 43.30%, and missing required effects in 25.36%; these categories overlap.
The release reports an overall pass@1 score of 67.16% for Claude Opus 5.5 and 57.37% for Kimi-K3, which it identifies as the strongest open-weight model in its table. The authors say Kimi-K3 scored within one point of GPT-6 Astra. The supplied material does not include the full results table or uncertainty estimates, so the rankings and score differences should be read as the authors’ reported results for this benchmark.
Why Backend State Changes Agent Evaluation
The benchmark targets a practical weakness in evaluating AI agents: a plausible response is not proof that a task was completed. An agent could tell a customer that a case is resolved while leaving the underlying ticket in the wrong status, entering an incorrect value, or making an unintended change. ThinkingBox’s executable checks are designed to reveal those outcomes.
That distinction matters for organizations considering agents for workflows such as refunds, support tickets, insurance claims and bookings. A mistake in a system of record can affect customer service or later business decisions even when the conversation appears successful. Repeated runs also offer a view of variability: an agent that passes once may not pass consistently.
The scores can help developers compare systems and locate workflow failures, but they do not establish how an agent will perform in any particular company’s systems. The benchmark is evidence about performance in its tested setup, not a guarantee of safe or reliable operation elsewhere.
As an affiliate, we earn on qualifying purchases.
From Tool Calls to Verified Outcomes
The ThinkingBox authors frame the work around the difference between what an agent does and what the system records. Tool-call validity and response quality are indirect measures: a tool can run without an error while the resulting record still fails the task’s requirements. The benchmark checks the terminal backend state and side effects directly.
One example in the release concerns a delayed $745 appliance order. The agent investigates the delay, opens a support ticket and records a timeline. The customer does not qualify for late-delivery compensation under the policy the agent checked, but a carrier exception remains open and the task requires the ticket to stay on hold. The agent instead marks it solved and replies without answering the customer’s underlying question. The executable check fails because the ticket status is solved rather than hold.
The release describes several measures: pass@1 is the share of attempts that succeed; pass@20 records whether a task succeeds at least once across 20 runs; and observed 20/20 counts tasks that pass every recorded run. These measures answer different questions and should not be treated as interchangeable evidence of reliability.
“A tool call is not an outcome.”
— Microsoft and Hugging Face, in the ThinkingBox release
business process automation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Reported Benchmark Results
The supplied release material does not state when ThinkingBox became available or provide all details of the evaluation, including the complete task specifications, model configurations and uncertainty estimates for the reported scores. It also does not establish whether passing 20 observed runs predicts long-term performance.
The reported failure rates and model scores come from the authors’ tested setup. The material does not show how the same agents would perform in every live business system, where records, integrations, policies and customer requests may differ. No independent replication or future milestone is identified in the supplied source.
backend system state monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
How Teams Can Test Their Own Workflows
The release says developers can run ThinkingBox through OpenEnv, using isolated MCP tool sessions to inspect tasks and compare agent results against executable checks. The supplied material does not give a timetable for further releases or independent evaluations.
For organizations weighing agent use, the practical next step is to test workflows that resemble their own operations and review both successful and failed runs. Teams will need to define the required records and side effects for each task; benchmark results alone cannot show whether an agent meets a company’s policies or performs reliably under its operating conditions.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is ThinkingBox?
ThinkingBox is an AI-agent benchmark released by Microsoft and Hugging Face. It checks whether agents leave business systems in the required state after carrying out workflows.
How does the benchmark judge an agent’s work?
It runs agents in isolated sessions using MCP tools, then checks backend records and side effects against executable task requirements. A valid tool call or fluent final response does not by itself count as a pass.
What do the reported failure figures mean?
The authors say 79,853 of 121,680 valid trials across 12 models failed executable checks. Their reported categories—wrong field values, unintended extra effects and missing required effects—overlap, so they should not be added together as separate portions of all failures.
Do the benchmark scores predict performance in a company’s systems?
Not by themselves. The reported scores describe results in ThinkingBox’s tested setup. The supplied material does not establish performance across other systems or show that passing 20 runs predicts long-term reliability.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
