arXiv:2608.19741nvidia/nemotron-3.5-lightning-30b-a3bAugust 20, 2026

One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business WorkflowsExplained for Beginners

Zhuochun Li, Youngmin Ko, Ali Keramati +9 more

Computation and LanguageDatabases

Abstract

Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call: agents must gather missing information over multiple turns, follow domain policies, coordinate dependent tools, and realize the correct persistent state transition without collateral effects. In this paper, we introduce Thinkingbox, a sandbox for tool-agent-user interaction that provides isolated MCP-compatible tool sessions, complete execution traces, and outcome evaluation over terminal backend state. Built on this sandbox, Thinkingbox-bench contains 507 policy-conditioned workflows across numerous scenarios, including retail, hospitality, auto insurance, neobank internal IT, and consulting IT/HR support. Each attempt is evaluated by task-specific executable checks that accept valid trajectories while rejecting wrong, missing, or extra effects; designated tasks additionally check required properties of the final response. Across proprietary and open-weight models, the strongest achieves 65.36% pass@1, but only 25.25% pass^20. Moreover, many failed trials show clean termination and valid state-changing actions, showing that response or tool-call-level signals are not clear proxies for end-to-end task completion. Thinkingbox-bench reveals a large gap between occasionally finding a successful trajectory and reliably completing stateful business tasks. We release both Thinkingbox and Thinkingbox-Bench: https://github.com/microsoft/thinkingbox

The Problem: Why “One Success” Isn’t “Reliability”

Imagine you’re a customer trying to change a hotel reservation over the phone. You explain your new dates, the agent pulls up your booking, and then—nothing happens. Maybe they put you on hold, maybe they type furiously, but the reservation stays exactly as it was. Or, worse, they accidentally change someone else’s booking while “helping” you.

In the world of AI agents, this scenario plays out every day. Recent benchmarks have gotten very good at measuring if an agent can produce a plausible answer or make a single correct tool call. But real-world business work is rarely a one-step operation. Agents must gather missing information over multiple turns, follow strict domain policies, coordinate dependent tools, and—most importantly—ensure the backend state ends up exactly right, without unintended side effects.

This paper introduces Thinkingbox and Thinkingbox-bench to answer a critical question: Are our AI agents actually reliably completing these stateful business tasks, or do they just occasionally stumble onto a success?

The gap is stark. The paper reveals that even the strongest models succeed on only about 65% of single attempts (pass@1), but when given 20 tries, they only succeed on all 20 attempts about 25% of the time (pass@20). More importantly, many "failures" look clean from the outside—the agent terminates the conversation politely, and the tool calls were valid—but the backend state is wrong. This means we’ve been measuring the wrong things. We’ve been rewarding smooth conversation, not reliable work completion.

How It Works: The Thinkingbox Sandbox

To measure this reliably, the authors built Thinkingbox, a sandbox that acts as a controlled universe for AI agents. Think of it as a highly sophisticated flight simulator for business workflows.

Instead of just asking an AI a question, Thinkingbox sets up a complete, isolated environment:

  • The Backend: A database representing a business (like a bank account, an insurance claim, or a hotel booking).
  • The Tools: MCP-compatible APIs the agent can call to interact with that database.
  • The User: A simulated person who might ask follow-up questions or provide missing info.
  • The Judges: Executable checkers that look at the final state of the database and the conversation to see if the task was truly completed.

The paper uses a concrete analogy that product managers and engineers will instantly recognize: Version Control (Git). Just as a developer might make a series of commits that look correct but ultimately break the build or merge the wrong code, an AI agent might make a series of tool calls that look correct but end up with the wrong database state. Thinkingbox ensures that to "pass," the final database state—and only the final state—must match the goal, regardless of the specific path the agent took to get there.

Key Results & Benchmarks: The Discovery vs. Reliability Gap

The benchmark, Thinkingbox-bench, contains 507 real-world workflows across five domains: retail/e-commerce, travel/hospitality, auto insurance, neobank internal IT, and consulting IT/HR. The results are humbling and eye-opening.

The Numbers:

  • Pass@1 (Single Try): The strongest model, GPT-5.4, achieved 65.36%.
  • Pass@20 (20 Tries): Even with 20 attempts, the pass rate dropped to 25.25%.
  • The "All-20" Success: Only about a quarter of tasks were reliably solved across 20 attempts.

The "Clean Failure" Phenomenon: Perhaps the most striking finding is that many failed trials show "clean termination." The agent finishes the conversation without error messages, and it does make valid state-changing actions. But often, it updates the wrong record or misses a required step. This proves that traditional metrics—looking at whether the model said the right thing or called the right API—are terrible proxies for actual task completion.

For example, in the auto insurance domain, models averaged only about 23% pass@1. In neobank support, it was roughly 30%. These aren't "hard" domains because the concepts are complex; they are hard because they require precise, multi-step coordination with zero tolerance for error.

Why It Matters: Key Takeaways

This research changes how we should think about AI agent reliability. Here are the four key takeaways:

  1. Fluent Tool Use ≠ Correct Workflow: A model can be great at generating tool calls (fluent) but terrible at the sequence of those calls (reliable). We need to evaluate the outcome, not the performance.
  2. The Retry Gap is Real: A high pass@1 score is misleading. If a model has a 65% pass@1, it means roughly one-third of the time, it will fail spectacularly or produce incorrect state changes. Reliability requires a much higher bar.
  3. Domain Specificity is King: General "smartness" doesn't transfer. A model might crush it on retail refunds but crash on auto insurance claims. Engineering teams need to benchmark their specific use cases, not just general capabilities.
  4. Side Effects are the Silent Killers: The most dangerous failures are the ones where the agent "succeeds" in changing a state, but changes the wrong one. In a business context, updating the wrong customer's insurance claim or accessing the wrong employee's records is a compliance and security nightmare.

Summary

Thinkingbox and Thinkingbox-bench provide a much-needed reality check. They move the evaluation goalpost from "can the agent talk a good game?" to "will the agent reliably get the business outcome right, every time?" The finding that pass@20 is roughly half of pass@1 suggests we have a long way to go before AI agents are trustworthy enough for high-stakes business workflows without constant human oversight.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →