Microsoft and Hugging Face benchmark AI agent reliability on database state
A new benchmark reveals that AI agents often report success while leaving incorrect database records, highlighting a gap between tool calls and actual outcomes.
Microsoft and Hugging Face have released ThinkingBox, a new benchmark that evaluates AI agents based on the actual database states they leave behind rather than just their generated text or tool calls. Published in October 2026, this joint effort shifts the focus from linguistic fluency to operational correctness in stateful business workflows.
What happened
The collaboration introduces a rigorous testing framework where AI agents are tasked with completing specific business workflows, such as processing refunds or updating customer tickets. Instead of grading the agent on whether it called the right tools or wrote a polite response, ThinkingBox inspects the terminal backend state. It checks if the database reflects the correct outcome after the agent finishes its work. This approach exposes a critical disconnect: an agent can execute all steps correctly according to its own logic but still fail to update the necessary records accurately.
To ensure robustness, the benchmark runs each of the 507 tasks twenty times against various large language models. This repetition highlights consistency issues that single-run tests miss. The results show that many models perform well on their first attempt but struggle to reproduce that success reliably. By focusing on executable checks of database fields, the authors provide a clearer picture of which models can be trusted for autonomous operations in production environments.
Key details
- ThinkingBox evaluates agents across 507 stateful business workflows, running each task 20 independent times.
- In a study of 121,680 valid trials, 67.24% of failures occurred even though the agent terminated cleanly and reported no errors.
- Claude Opus 5.5 achieved the highest overall pass@1 score at 67.16%, followed closely by Claude Opus 5 at 66.50%.
- Kimi-K3 demonstrated the broadest capability, solving 93.89% of tasks at least once, but only maintained consistency on 13.41% of tasks across all 20 attempts.
- Only three models retained most of their single-attempt performance over 20 repeats: GPT-6 Astra (78%), Claude Opus 5.5 (71%), and Claude Opus 5 (71%).
- The benchmark is available via OpenEnv, allowing developers to test models against isolated MCP tool sessions.
Background
Traditional AI benchmarks often rely on static datasets or evaluate the quality of natural language responses. These methods assume that if an agent sounds correct and uses the right tools, the job is done. However, in software systems, the ultimate truth lies in the data store. If a customer service agent says a ticket is resolved but the database status remains open, the workflow has failed regardless of the agent's confidence.
ThinkingBox addresses this by treating every agent trajectory as a claim and the database state as evidence. It uses isolated environments to prevent side effects from contaminating other tests. This method aligns more closely with how engineering teams verify software integrity, focusing on idempotency and state correctness rather than just functional completion. It moves beyond "does it work?" to "does it work every time?"
Why it matters
For teams integrating AI agents into their internal tools or customer-facing products, this benchmark offers a reality check. Relying on single-run success metrics can lead to fragile automations that break under slight variations in input or model behavior. The high rate of silent failures—where the agent reports success but the data is wrong—poses significant risks for financial transactions, inventory management, and user account updates.
Understanding the difference between breadth and consistency helps in model selection. A model like Kimi-K3 might be suitable for exploratory tasks where coverage is key, while Claude Opus 5 is better suited for repetitive, high-stakes operations where reliability is non-negotiable. Teams can use these insights to design fallback mechanisms and human-in-the-loop checkpoints for tasks where model consistency drops below acceptable thresholds.
What you can do
- Evaluate your current AI agents using the ThinkingBox methodology by checking final database states rather than just log outputs.
- Run critical workflows multiple times in staging environments to measure consistency before deploying to production.
- Prioritize models with high observed 20/20 scores for tasks involving financial data or irreversible actions.
- Implement automated verification scripts that query the database after agent execution to confirm expected state changes.
- Use isolated test environments to prevent cross-contamination when benchmarking new agent configurations.
- Review agent logs for silent failures where tool calls succeed but required field updates are missing or incorrect.


