The Agent Said It Was Done. The Database Knew Better.
A new benchmark from Microsoft and Hugging Face runs AI agents through real business tasks 20 times each, then checks the records they leave behind. Most models fail far more often than a single test suggests.

Key points
- ThinkingBox, a new AI agent benchmark built by Microsoft researchers and released through Hugging Face, tests 507 real business workflows and checks the actual database state an agent leaves behind, not just whether it gave a polite answer.
- Claude Opus 5.5 scored highest overall at 67.16% on a single attempt, but retained only 71% of that score when each task ran 20 times in a row.
- GPT-6 Astra was the most consistent model tested, keeping 78% of its single-attempt score across all 20 repetitions.
- In 121,680 test runs, 67.24% of failed attempts still ended cleanly with no obvious error, meaning the agent appeared to succeed while the database said otherwise.
- Kimi-K3, an open-weight model that runs on your own hardware without a subscription, solved 93.89% of the benchmark at least once, the broadest coverage of any model tested.
A customer's $745 kitchen appliance has been stuck at a courier depot for fifteen days. An AI agent handles her complaint. Nine tool calls: it reads the refund policy correctly, opens a support ticket, closes it with a cheerful "Is there anything else I can help you with?"
The carrier exception is still open. She never got a real answer. The database says the ticket is "solved" when it should be "hold."
An AI checking the agent's actions would see nine well-formed steps and declare success. The database disagrees.
That gap is exactly what ThinkingBox measures. Developed by researchers at Microsoft and released through Hugging Face, the benchmark runs AI agents (software that carries out multi-step tasks on its own) through 507 business scenarios drawn from retail, travel, auto insurance, neobank and consulting support. It grades the actual state of a backend database after the agent finishes, not the text of its final reply.
Why does one successful run not mean much?
Real software runs the same task thousands of times, not once. ThinkingBox runs every task 20 times from a clean starting state, then asks three questions: how often does the agent succeed on a random try, can it solve this at all, and does it get it right every single time.
The results are humbling.
| Model | Overall pass@1 | % of score kept across 20 runs |
|---|---|---|
| Claude Opus 5.5 | 67.16% | 71% |
| Claude Opus 5 | 66.50% | 71% |
| GPT-5.4 | 65.36% | not reported |
| GPT-6 Astra | 58.31% | 78% |
| Kimi-K3 (open-weight) | 57.37% | ~8% |
| DeepSeek-V4-Pro | 43.26% | ~8% |
GPT-6 Astra is the most consistent: it holds 78% of its single-attempt performance across all 20 runs. Claude Opus 5.5 and Claude Opus 5 each keep 71%. At the other extreme, Kimi-K3, GLM-5.1 and DeepSeek-V4-Pro each retain roughly 8%, so a score that looks reasonable on a leaderboard collapses almost entirely under repetition.
Kimi-K3 still earns a notable distinction: it solved 476 of 507 tasks at least once, the broadest coverage in the field. Breadth and reliability aren't the same thing.
What does this mean for anyone using AI tools at work?
Simple version: an AI assistant that completes a task correctly in a demo may not do so reliably in daily use. Two-thirds of failed runs in this benchmark ended with no visible error. The tool seemed fine; the actual record was wrong.
For businesses building AI agents into customer service or internal IT, it's worse than that. Our 2 October story on ServiceNow's AutoSynthData found the strongest model tested cleared barely a third of realistic enterprise tasks. These figures show one reason why: a model that handles retail complaints at 68% accuracy can drop to 8% on auto insurance claims from the very same company.
The benchmark's available to run yourself through the OpenEnv platform linked in the ThinkingBox paper on arxiv.org. No subscription required.
The lesson the industry keeps relearning: a confident answer and a correct outcome aren't the same thing. The only way to know which you got is to check what the system actually wrote down.
Common questions
Does this affect consumer AI tools like Claude or ChatGPT?
Not directly. ThinkingBox tests AI agents built into business software, not the chat apps most people use day-to-day. But the same models power both, so consistency problems in a benchmark are a real signal about how those models behave under pressure.
What is an "open-weight" model and why does it matter here?
An open-weight model is one whose underlying code is publicly released, so a company or developer can run it on their own computers without a subscription fee. Kimi-K3 showed the widest task coverage of any model in the benchmark, which matters for organisations looking to cut costs, though its consistency score is far weaker than the top proprietary models.
Where can I read the full results?
The complete paper, including all benchmark tasks and failure traces, is available at arxiv.org.



