Skip to content
AI.info

The Pulse

Microsoft and Hugging Face Test Agents Across 507 Workflows

Microsoft and Hugging Face released ThinkingBox, a benchmark that tests AI agents on 507 business workflows and checks the records they change. Its 20-run evaluations found that many agents can finish cleanly while leaving required work und

Microsoft and Hugging Face Test Agents Across 507 Workflows

AI.info Team ·

67% of failed runs still looked finished

In a test of AI agents across business workflows, 67.24% of failed attempts ended without a tool error—even though the systems had not completed the required work. Microsoft and Hugging Face published the results on October 3, 2026, alongside ThinkingBox, a sandbox and benchmark built to check what agents actually change in backend records.

The figure comes from an analysis of 121,680 trials across 12 models. Of the attempts that failed the benchmark’s executable checks, many had made state-changing tool calls and then stopped normally. The checks still found wrong field values in 77.61% of those failures, unintended extra effects in 43.30%, and missing required effects in 25.36%; those categories overlap.

The finding challenges a common shortcut in evaluating tool-using systems: treating a fluent final answer or valid tool call as evidence that a task is complete. “Thinkingbox-bench reveals a large gap between occasionally finding a successful trajectory and reliably completing stateful business tasks,” write Zhuochun Li et al. in the ThinkingBox paper.

ThinkingBox checks the records, not just the transcript

ThinkingBox-Bench contains 507 policy-conditioned workflows covering retail, hospitality, auto insurance, neobank support, and consulting IT and HR. Each task gives an agent a goal, an initial backend state, available tools, and applicable policies. The system then checks the resulting records and side effects against executable requirements.

Every task runs 20 times from a freshly initialized backend, with attempts isolated from one another. The benchmark reports several distinct measures: pass@1 for the single-attempt success rate, pass@20 for tasks solved at least once in 20 attempts, and observed 20/20 for tasks that pass every recorded attempt. That last measure addresses repeatability directly; passing once does not show that an agent can reliably handle the same workflow.

The public tasks are synthetic reconstructions of enterprise patterns, not records from real customers. In the announcement’s example, an agent handles a delayed appliance delivery, opens a support ticket, then incorrectly marks it resolved while the carrier exception remains open. Its tool calls look orderly; the ticket’s final status reveals the mistake.

Models trade breadth for consistency

The results separate occasional capability from dependable execution. Kimi-K3 solved at least one attempt on 476 of the 507 tasks, or 93.89%, but passed all 20 attempts on only 68 tasks. Claude Opus 5 solved at least one attempt on fewer tasks, while completing every attempt on 47.53% of the benchmark.

Claude Opus 5.5 achieved a higher single-attempt score than Opus 5, 67.16% versus 66.50%, but both passed exactly 241 tasks on all 20 attempts. The figures show why the benchmark distinguishes getting a task right once from getting it right consistently. Results also varied by workflow: the announcement reports a large gap between retail and auto-insurance performance for some models.

A test environment, not proof of production reliability

ThinkingBox isolates each run in an MCP-compatible tool session and compares the final backend state with task-specific checks. Of the 507 tasks, 477 rely on state checks alone; the remaining 30 also use response rubrics for requirements that have no straightforward database value. The model receives the task, dialogue, and tool schemas, while the benchmark keeps its expected state and grading details outside the agent’s view.

The framework and benchmark are available through Hugging Face and OpenEnv, with the executable benchmark released as version 1.0. The results offer developers a way to inspect failures and repeat the same workflow under controlled conditions, rather than relying only on an agent’s own account of what it did.

But the release does not establish how closely the synthetic workflows predict failure rates in live enterprise systems. Whether those repeated-run gaps carry over to real records remains an open question.

Sources

Explore

More articles