AI & LLMs

Grading AI agents on database state instead of text output

Microsoft and Hugging Face released ThinkingBox, a framework that evaluates AI agents by verifying their final database records rather than their conversational logs.

Illustration of verifying database records against AI agent output
Illustration created for this article

On October 3, 2026, Microsoft and Hugging Face jointly published ThinkingBox, a new evaluation framework for autonomous AI agents. This approach shifts the focus from grading the natural language responses an agent generates to verifying the actual state changes it leaves in backend systems. The initiative highlights a critical gap in current testing methods where agents can appear successful in logs while failing to execute required business logic correctly.

What happened

The core insight behind ThinkingBox is that an agent can perform a series of correct tool calls, generate a polite and professional closing message, and still fail to achieve the desired outcome. In a provided walkthrough, a support agent handled a case involving a delayed kitchen appliance. The agent read the policy, made nine tool calls, and closed the ticket with a message stating the issue was resolved. However, the required end state for such a case was to place the order on hold, not mark it as solved. While a traditional evaluator looking at the conversation transcript would mark this as a success, a check against the database reveals the customer never received the correct resolution.

This discrepancy is not merely theoretical. The authors analyzed a common-set ablation involving 121,680 valid trials across 12 different models. They found that 79,853 attempts failed executable checks. Among these failures, 67.24% terminated cleanly, invoked state-changing tools, and reported no errors. Despite the clean exit, further inspection revealed wrong field values in 77.61% of these cases, unintended extra effects in 43.30%, and missing required effects in 25.36%. These overlapping issues demonstrate that a green trace in monitoring tools does not guarantee a passed state in the system of record.

To address this, the framework proposes a strict validation loop. Teams must define the required end state in specific fields before the run begins. After the agent performs its last mutating action, the system must read the record back directly from the database, ignoring the agent's summary. Customer-facing actions should only proceed if this readback confirms the fields match the requirements. This process creates a receipt that documents exactly what matched, what did not, and any unexpected writes, ensuring that the definition of done is based on data integrity rather than linguistic fluency.

Key details

  • ThinkingBox was published on October 3, 2026, by Microsoft and Hugging Face.
  • In 121,680 trials, 67.24% of failed attempts still terminated cleanly without reporting errors.
  • Wrong field values were found in 77.61% of the cleanly terminated but failed attempts.
  • Claude Opus 5.5 achieved a pass@1 score of 67.16% but only passed all 20 attempts on 241 tasks.
  • Kimi-K3 solved 476 of 507 tasks at least once but only achieved consistent success on 68 tasks.
  • Approximately 79.9% of failures were attributed to tool handling issues rather than reasoning errors.

Background

In traditional software testing, engineers often rely on unit tests that check function outputs or integration tests that verify API responses. For AI agents, which interact with systems through a sequence of tool calls, evaluations have largely focused on the "trajectory" or the path the agent took. This includes checking if the right tools were called in the right order and if the final message was appropriate. This method assumes that if the steps look correct, the outcome is correct. However, autonomous agents operate in non-deterministic environments where tool responses can vary, and internal system states may not align with the agent's perception.

ThinkingBox introduces the concept that a trajectory is merely a claim, while the database state is the evidence. By treating the agent's actions as a hypothesis about the system state, developers can use post-execution verification to validate that hypothesis. This moves beyond simple pass/fail metrics based on single runs. It emphasizes repeatability, asking whether an agent can achieve the correct state consistently across multiple attempts from a clean starting point, rather than just getting lucky once.

Why it matters

For teams running self-hosted software or managing internal automation, this distinction is vital for reliability. If you deploy an agent to handle user support, inventory updates, or financial transactions, trusting the agent's final message is a significant risk. An agent might confidently report that a refund was processed because the payment gateway returned a generic success code, even if the internal ledger was not updated due to a race condition or a schema mismatch. Without verifying the actual record, your team may remain unaware of systemic data corruption until customers complain.

Furthermore, this approach impacts how you select and optimize models. The benchmark data shows that higher headline scores do not necessarily translate to better dependability. Claude Opus 5.5 had a slightly higher pass@1 score than Claude Opus 5, but both models achieved consistent success on the same number of tasks. Choosing a model based on its ability to solve a task at least once, rather than every time, leads to fragile production systems. Understanding that most failures stem from tool handling rather than reasoning suggests that engineering efforts should focus on robust retry policies and smaller, more reliable tool surfaces instead of simply upgrading to larger models.

What you can do

  • Define the required end state for your agent workflows using specific database fields before execution.
  • Implement a post-write readback step that queries the system of record after the agent completes its task.
  • Block customer-visible notifications until the readback confirms that all required fields match the expected state.
  • Generate a diff receipt for each run that lists matched fields, mismatched fields, and any unintended side effects.
  • Run your critical workflows multiple times from a clean state to measure consistency, aiming for every-of-k success rather than best-of-k.
  • Audit your current evaluation metrics to ensure they are not rewarding agents for polite but incorrect closures.

More news

All news