THE SIGNAL IN ONE SENTENCE
An AI agent can make nine valid tool calls, open the right support ticket, quote the right policy and still leave the customer in the wrong state. That is the useful annoyance at the center of ThinkingBox, a newly released sandbox and benchmark from Microsoft's Copilot Studio team and research collaborators. Instead of grading an agent by its final answer or by whether its tool calls looked sensible, ThinkingBox inspects the backend after the work is over. Did the refund really happen? Is the ticket on hold instead of marked solved? Was the correct record changed? Did the agent avoid creating an extra side effect? The database gets the last word. The benchmark contains 507 synthetic business workflows across retail, auto insurance, travel, neobank operations and consulting support. Every task starts from a clean, isolated backend state. A simulated user holds some private facts and supplies them only when the agent asks. The agent gets a policy and a set of tools. At the end, executable checks compare the actual records with the required outcome. Then the whole thing happens nineteen more times. That repetition is not decorative. Agent behavior is probabilistic. A model that handles a cancellation correctly on Monday and mangles the same case on Tuesday is not a reliable cancellation system, even if its average demo looks smooth. The plain signal is simple: a completed tool call is not a completed job, and one successful run is not a dependable workflow. ThinkingBox was introduced in a paper first submitted in August and revised on October 1. Microsoft and Hugging Face published the runnable OpenEnv version and a detailed results article on October 3. The framework code is MIT licensed, the benchmark data uses CDLA Permissive 2.0, and the OpenEnv adapter uses the BSD 3-Clause license. The release matters because it moves a familiar complaint into executable form. People have watched agents say "done" after editing the wrong row, skipping a required field or failing to recover from a tool error. Teams often catch those mistakes by reading traces after something goes wrong. ThinkingBox makes the final state itself the test target. One benchmark example begins with a customer whose kitchen appliance is fifteen days late and stuck in a carrier exception. The agent checks the order, tracking, customer profile and refund policy. It opens a ticket and correctly determines that the customer does not qualify for late-delivery compensation. Then it marks the ticket solved. The shipment problem is still open, so the ticket should be on hold. The agent also fails to give the customer a useful answer. A grader focused on syntax would see nine valid tool calls. A grader focused on the final database sees the wrong status. That distinction becomes uncomfortable at scale. In a common-set analysis covering 121,680 valid trials across twelve models, the researchers report 79,853 failed attempts. Among those failures, 67.24 percent still ended cleanly, used a state-changing tool and showed no final tool error. The executable checks found wrong field values in 77.61 percent of failed attempts, unintended extra effects in 43.30 percent and missing required effects in 25.36 percent. Those categories overlap, so they should not be added together. In other words, many failures did not look like crashes. They looked like finished work. That is the dangerous shape of an agent error. A broken API call is noisy. A wrong database update followed by a polished summary is quiet. The benchmark reports three views of performance. Pass at one is the share of individual attempts that succeed. Pass at twenty asks whether a task succeeded at least once across twenty tries. ThinkingBox also reports the literal number of tasks that passed all twenty recorded attempts. Those questions are not interchangeable. A model can solve many kinds of tasks at least once and still be erratic on nearly all of them. The strongest open-weight model in the published results, Kimi-K3, solved 476 of 507 tasks at least once. Yet only 68 tasks passed in all twenty attempts. Claude Opus 5 showed the opposite tradeoff. It solved fewer tasks at least once, but 241 tasks passed all twenty attempts. Claude Opus 5.5 had a slightly higher single-attempt score and broader coverage, yet it also reached exactly 241 perfect-repeat tasks. That result punctures an easy purchasing story. A newer model can improve its average score without expanding the set of jobs that it performs dependably. The benchmark's headline table puts Claude Opus 5.5 first on overall pass at one with 67.16 percent. Claude Opus 5 follows at 66.50 percent, GPT-5.4 at 65.36 percent and GPT-5.6 Sol at 61.91 percent. Kimi-K3 leads the listed open-weight models at 57.37 percent. Those are results inside this benchmark, not universal model rankings. Domain variation is large. The authors report Claude Opus 4.6 at 68.62 percent on retail workflows and 8.30 percent on auto insurance. Across the listed models, retail averaged 59.52 percent while auto insurance averaged 33.83 percent. A team choosing an agent should therefore resist the temptation to buy the top row and call the problem solved. The right question is whether the agent survives repeated tests built from the team's own policies, tools, records and failure costs. The failure breakdown offers a practical clue. ThinkingBox assigns one deterministic diagnostic signature to each failed trace. In the authors' ablation, 79.9 percent of failures were categorized as tool usage, 10.3 percent as wrong state updates, 7.0 percent as incomplete user resolutions and 2.9 percent as no state-changing action. These are observable labels averaged across models, not a proof of a unique underlying cause. Still, the pattern is helpful. Many agents get far enough to try the job and then mishandle a failed precondition, empty lookup, tool error or recovery path. That suggests a boring but valuable engineering response. Before committing a consequential change, read the record back. Compare the final state with explicit invariants. Classify tool errors so retry logic knows which failures are safe to repeat. Limit the tool surface to what the workflow actually needs. Require a person to approve changes that are expensive or impossible to reverse. This does not require every company to reproduce all 507 tasks. A support team could begin with ten representative workflows: a refund, a delayed shipment, an address change, a duplicate ticket, a cancellation, a policy exception, a failed payment, a sensitive-data request, an escalation and a case where the correct action is to do nothing. Seed a clean test database, run each case repeatedly and check the resulting state with code. The final answer can be tested too, but it should not substitute for the records. ThinkingBox grades 477 of its 507 tasks on state alone. Thirty tasks add a narrow response rubric for requirements that do not have a clean database value, such as whether the agent disclosed an uncertainty. That split is sensible. Some obligations live in language, but many operational claims can be checked directly. The trust boundary also matters. The model sees the task, dialogue and tool schemas. It does not see the golden state, assertions, evaluator internals or credentials. Every attempt receives a fresh isolated MCP session, so two runs cannot accidentally share a row or cached tool result. A side-effect extractor identifies what changed, and deterministic judges reject wrong, missing or extra effects. This design is much stronger than asking another language model whether the first model sounded correct. It is not magic. The workflows are synthetic reconstructions modeled on enterprise patterns. The customers are not real. A benchmark cannot capture every ugly property of a production system: stale caches, partial outages, conflicting permissions, race conditions, hidden dependencies, unusual policy exceptions or the social pressure that makes a human operator override a warning. The model endpoints, prompts, tool implementations and evaluation harness also shape the result. A different agent scaffold may perform differently with the same underlying model. The authors pin framework code, benchmark data and bundle hashes to make canonical results reproducible, but adopters still need to test their own complete system. Cost estimates need equal care. The article prices recorded token usage at undiscounted list rates from a September 20 OpenRouter snapshot, with separate handling for Claude Opus 5.5 pricing. The authors call the result a comparative efficiency index, not a production invoice. That caveat matters because the benchmark offers a clever new unit: cost per dependable task. It divides the cost of all twenty runs by the number of tasks that passed all twenty times. GPT-5.4 had the lowest reported cost per dependable task at $6.80, but only 128 tasks met the perfect-repeat bar. GPT-6 Astra reached 231 dependable tasks at $7.45 each. Claude Opus 5.5 reached 241 at $7.80. Claude Opus 5 also reached 241, but at $13.30. The cheapest single success was not always the cheapest dependable workflow. That is a useful budget conversation. A low per-call price can look excellent until retries, review, repair and customer recovery are included. If a workflow changes money, coverage, access or identity records, the cost of one quiet mistake may dwarf the model bill. ThinkingBox does not prove that an agent should run a bank, insurer or support desk without people. It shows how to make a narrower claim with better evidence: under these policies, tools and starting states, this system produced the required end state this often, and it repeated that result this many times. That sentence is less thrilling than an autonomous-agent demo. It is also much closer to what an operations manager needs. The release leaves several important questions open. First, the authors have not reported how much specific interventions improve the benchmark. Read-back verification, bounded retries, smaller tool sets and approval gates are sensible recommendations, but their lift and failure costs still need to be measured. Second, a perfect twenty-run record is evidence, not a guarantee. A rare failure can hide beyond the twentieth attempt. The right repeat count depends on risk. A newsletter draft and an insurance cancellation should not share the same confidence threshold. Third, deterministic state checks are only as good as the specification. If the test encodes the wrong policy, the agent can pass perfectly while doing the wrong thing in the real world. Someone must review the invariants, exception rules and forbidden side effects. Fourth, the benchmark mainly evaluates bounded workflows. Open-ended research, negotiation, care work and creative tasks do not always end in a database state that can be declared correct. Outcome verification is powerful where the outcome is legible. It should not be stretched into a fake answer for work that remains ambiguous. The most practical lesson is not to chase a universal reliability score. Build a small evidence loop around every consequential agent action. Define what must be true before the action. Define what must be true afterward. Record the exact side effects. Read the state back from the system of record. Run the same case enough times to expose variance. Keep the evaluator outside the agent's control. Route ambiguous or irreversible cases to a person. Then publish both the ordinary success rate and the repeat rate. An agent's sentence is a claim about the work. The records are the evidence. And if the records disagree, believe the database before the chatbot.
01
WHAT ACTUALLY CHANGED
Microsoft and Hugging Face released ThinkingBox through the OpenEnv interface on October 3, with a runnable harness, benchmark data and installation guidance.
ThinkingBox-Bench contains 507 synthetic, policy-conditioned workflows across retail, auto insurance, travel, neobank operations and consulting support.
Every task is run twenty times from an isolated clean backend state, and executable checks inspect the final database state and side effects.
In a 121,680-trial common-set analysis across twelve models, 79,853 attempts failed the executable checks even though many ended without a visible tool error.
The release reports ordinary single-attempt success, at-least-once breadth and the literal number of tasks that passed all twenty recorded attempts.
The framework code, benchmark data and OpenEnv adapter are published under permissive licenses.
02
WHY THIS MATTERS
A valid tool call proves that an operation was attempted, not that the correct business outcome was reached.
Quiet failures can leave wrong fields, missing effects or unintended extra changes while the agent presents a polished completion message.
Repeated evaluation exposes a gap between occasional capability and operational dependability.
Database-state checks give teams an objective target for workflows that affect money, records, access, claims or customer obligations.
Model rankings can change sharply by domain, so teams need tests built from their own policies and tool surfaces.
Cost per successful call can hide the expense of retries, review, repair and recovery after inconsistent behavior.
03
WHERE IT COULD HELP
- Create executable before-and-after checks for refunds, cancellations, claims, bookings, access changes and support tickets.
- Read critical records back from the system of record before telling a user that the work is complete.
- Repeat representative tasks from clean state instead of relying on one successful demo trace.
- Track wrong, missing and extra side effects separately so fixes target the real failure shape.
- Classify tool errors and allow bounded retries only for failures that are safe to repeat.
- Reduce each workflow to the smallest tool set and permissions it actually needs.
- Keep golden state, assertions and evaluator credentials outside the agent context.
- Add a human approval gate for irreversible, high-value or legally consequential state changes.
- Report both ordinary success and perfect-repeat rates with the exact number of trials.
- Recheck costs after including retries, review time and the expected cost of a wrong outcome.
KEEP A HAND ON THE WHEEL
ThinkingBox evaluates 507 synthetic reconstructions of enterprise workflows, not live customers or a universal sample of business operations. Results depend on the model endpoint, scaffold, tools, policies, prompts and evaluator configuration used in the campaign. Twenty successful repeats do not prove that a twenty-first run cannot fail, and deterministic checks can faithfully enforce a mistaken specification. The cost analysis uses a September 20 price snapshot and is a comparative index rather than a production invoice. The authors recommend state verification, targeted retries, smaller tool surfaces and human approval for hard-to-reverse actions, but they have not yet reported the measured improvement from those interventions on this benchmark. Watch for independent reproductions, additional domains, tests with real production failure conditions, measured mitigation lift, stronger uncertainty estimates and evidence that teams can write correct executable policies before agents are allowed to change consequential records.
04
TERMS WORTH KEEPING
SOURCES AND VERIFICATION STATUS
This article was written from the materials below. Product claims and dates were checked against those sources on October 4, 2026.
THE PUBLICATION ENGINE
WANT A SIGNAL OF YOUR OWN?
We build source-grounded publications, private briefings, and editorial systems for organizations with something useful to say.
WORK WITH US