τ^τ-Bench: An Environment for End-To-End, Realistic Agent ConstructionExplained for Beginners
Quan Shi, Keshav Dhandhania, Karthik Narasimhan +1 more
Abstract
LLM agents are rapidly becoming production software, deployed to handle customer service, adjudicate disputes, and operate internal systems. Notably, the work of building them is increasingly handed to coding agents, yet existing benchmarks say little about whether an AI system can deliver one under the conditions of a real client engagement. We introduce τ^τ-bench (pronounced hyper-tau-bench), a benchmark that makes agent construction the task. A developer agent is given the records a business actually keeps, a client who holds requirements, a production API that operations must run through, a codebase to inherit, and limits on serving cost and models: the same starting point a real engagement provides. From these it must deliver a complete customer-service agent, scored by deploying that agent against held-out simulated users. Across 53 tasks spanning four domains, the strongest configuration, Claude Opus 5 under Claude Code, passes just 23.9% of evaluation simulations. Meanwhile, an expert-authored reference ceiling scores 82.2%. The failures mirror ones human agent developers see: models issue shallow queries in place of deep comprehension of the records, communicate almost nothing to the client, and experiment too little with agent architecture and serving spend, shipping the first design that runs. We aim for τ^τ-bench to turn the work of cooperative agent building into a measurable target for coding agents.
τᵀ-Bench: Building Customer-Service Agents the Way the Real World Does
If you work in AI, you’ve likely noticed a peculiar gap. We have benchmarks that measure how well a finished agent chats with simulated users. We have benchmarks that measure how well a model writes code. But we have very little that measures whether an AI system can build a production-grade agent from scratch, under the messy, constraint-heavy conditions of an actual client engagement.
A new benchmark, τᵀ-bench (pronounced "hyper-tau-bench"), attempts to fill that gap. It doesn’t score an agent’s conversational skill; it scores a developer agent’s ability to construct one. The results, drawn from 53 realistic tasks across four domains, are a sobering reality check on where we stand.
Below is a breakdown of the problem, the mechanics, the results, and why this matters.
1. The Problem: The Gap Between "Can Chat" and "Can Build"
The paper identifies a fundamental mismatch in how we evaluate AI agents today.
The Real-World Engagements LLM agents are rapidly becoming production software. Enterprises deploy them to handle customer service, adjudicate disputes, and operate internal systems. However, building such an agent is difficult in two ways.
First, there is the problem of recovering the specification. The knowledge an agent needs is scattered: standard operating procedures, support transcripts, fee schedules in spreadsheets, and—critically—knowledge that lives only in human stakeholders. Those requirements surface under questioning and shift as the business changes.
Second, there is the problem of construction itself. It rarely begins from a blank slate or an unbounded budget. The developer may inherit an existing codebase whose behavior must be preserved, and must satisfy hard constraints on cost, latency, and available models.
The Benchmark's Premise τᵀ-bench flips the script. Instead of asking "Can this agent chat well?", it asks: "Given a business's records, a client with requirements, a production API, an inherited codebase, and a budget, can you deliver a complete customer-service agent?"
The developer agent must deliver a complete agent, which is then scored by deploying it against held-out simulated users. The starting point is exactly what a real engagement provides.
2. How It Works: The Technical Mechanics
The benchmark constructs each task from seven levers, combining to create a realistic development environment.
The Components a Developer Receives
For each of the 53 tasks, the developer agent is given:
- A Document Corpus (A): The records a business actually keeps—handbooks, support transcripts, images, audio, fee schedules, email threads. Across the four domains, the transformations produce 2,868 distinct evidence artifacts, totaling over 5.5 million tokens of text.
- A Simulated Human Client (C): Holds requirements that the records omit. The developer must interrogate this client over multi-turn conversation to uncover them.
- A Client-Operated REST API (T): The production API that operations must run through. The agent must integrate with this, designing its action space and writing adapters beneath it. Some tasks plant deterministic defects in the runtime.
- A Starting Implementation (π₀): Often, there is an existing codebase to inherit. These are genuine implementations with genuine defects—partial coverage, stale values, misread rules.
- A Menu of Models (M) and a Credit Budget (b): The constructed agent must run on a fixed menu of models within a budget on mean serving cost per conversation. This recreates the incentives real teams build under.
The Developer's Workflow
The developer works in a sandboxed environment with no internet and no reference implementation. The evaluation suite is withheld; feedback before submission must come from simulations the developer authors itself.
The developer works the levers in sequence:
- Recover the Specification: Read the records and ask the client.
- Build the Agent: Edit the inherited workspace, design tools, and write code.
- Self-Evaluate: Script a customer, assert the outcome, run local tests.
- Submit · Sealed Eval: Submit to held-out evaluation tasks; cost is counted.
An Analogy for Engineers
Think of τᵀ-bench like being handed a new project at a startup on your first day.
- The Records: A disorganized folder of wikis, old emails, and meeting notes (the corpus).
- The Client: The founder who knows the market but hasn't written the requirements down (the simulated client).
- The API: The company's internal service you must interact with, possibly with bugs (the REST API).
- The Budget: A limited amount of cloud credits you must spend wisely (the credit budget).
- The Deliverable: You must build a working prototype of the customer-service bot, complete with tools and logic, before the end of the week.
The benchmark measures whether you can navigate this scenario end-to-end.
3. Key Results & Benchmarks
The headline number is stark. Across 53 tasks spanning four domains (airline, retail, telecom, and banking), the strongest configuration—Claude Opus 5 under Claude Code—passes just 23.9% of evaluation simulations.
For context, an expert-authored reference ceiling scores 82.2%. The gap is massive: the best AI builder is less than a third as effective as an expert human.
Performance by Domain
The gap varies significantly by domain, highlighting where the challenge is hardest:
| Domain | Developer Pass Rate | Reference Ceiling | Gap |
|---|---|---|---|
| Airline | 55.9% | 82.2% | 26.3% |
| Retail | 72.8% | 82.2% | 9.4% |
| Telecom | 48.2% | 82.2% | 34.0% |
| Banking | 5.9% | 82.2% | 76.3% |
Banking is by far the most complex domain. While airline, retail, and telecom span 85, 119, and 155 atomic facts, banking's corpus carries 2,969 atomic facts. A single banking task can draw on up to 580 of them—two to seven times an entire other domain.
Translating the Numbers
When the paper says the strongest config passes 23.9%, it means that, on average, for every 100 simulated customer conversations the built agent was deployed against, it achieved the correct outcome in only about 24 of them.
The reference ceiling of 82.2% means that, even with perfect knowledge of the domain (which the developer doesn't have), an expert can build an agent that gets it right over 8 times out of 10.
Cost and Effort
The study also tracked the "friction" of building:
- Build Time: Averages 47.9 minutes under Codex, 216.3 under Claude Code, and 205.7 under Kimi Code.
- Build Cost: Developer token spend at API list prices ranged from 42 (Claude Code).
- Serving Budget Mismanagement: In all, 21 builds overshot the budget, and the penalty erased an otherwise positive score for ten of them. However, underusage was more common: on average, developers spend about half their serving budget and do not buy stronger models or longer context with the rest.
4. Why It Matters: Key Takeaways
The failures observed are not random; they mirror the specific pain points human agent developers face. Here are the four most significant takeaways:
1. Shallow Querying vs. Deep Reading
The developers read the corpus by keyword: each opened fewer than 80 of its roughly 1,700 files (in the banking domain), ran 22–53 corpus-wide text searches, and built from what came back. The agents they shipped simply do not know the policy; one answered, "I don't have a verified catalog of Rho-Bank's current credit-card offerings."
The Insight: For large natural language corpora, grep may be an effective way to index into a codebase, but it can miss a lot of crucial evidence. Determining when search heuristics are enough—and when it's necessary to actually read through files—is a judgment today's models make poorly.
2. The Client is Rarely Interviewed
On client-enabled tasks, where 20–25% of requirements live only with the simulated client, developers asked it at most four questions before shipping. Every question was triggered by a conflict or contradictory records. The misses were often avoidable; one developer wrote open questions into a planning file and never sent them.
The Insight: Builds that ask the client outscore builds that never do by roughly a factor of three. The client holds the "unwritten facts" that are essential for a complete agent.
3. Inherited Code is Largely Rewritten
Every seeded build audited opened the inherited code in its first twenty steps and read all of it. However, even when the seed carried an architecture known to lead to better scores, the developers threw it out. For example, one audit called a seed's intent routing—a design that doubled the telecom score when given as a hint—a defect, and every build replaced the starter tools and agent wholesale.
The Insight: Developers cannot reliably distinguish working code from broken code in a new codebase. The working parts often have to be rebuilt from the corpus.
4. Budget Mismanagement in Both Directions
Developers struggle with the serving budget in both directions. While 21 builds overshot the budget (triggering penalties), the far more common issue was underusage. On average, developers spend about half their serving budget and do not buy stronger models or longer context with the rest. Some submitted with half or more of the build clock left, and nineteen requirements still sat with the client, unasked.
The Insight: Under unit economics, the developer must work a cost–quality Pareto frontier: routing work across models, orchestrating cheaper ones, and debloating prompts and context, since every token is billed.
Conclusion
τᵀ-bench is a significant step toward making the work of cooperative agent building a measurable target for coding agents. It moves beyond evaluating an agent's ability to chat and instead evaluates the harder task of building one from real-world constraints.
The results show that while AI systems can build agents that run, they struggle to build agents ready for deployment. The strongest configuration passes less than a quarter of evaluation simulations, and the failures mirror the specific challenges human developers face: shallow queries, failing to interview the client, discarding working inherited code, and mismanaging the serving budget.
As the field moves toward more autonomous AI systems, benchmarks like τᵀ-bench will be essential for ensuring that these systems can not only operate in production but also be constructed reliably under the pressures of a real client engagement.
If you are building AI agents or evaluating LLM capabilities, this paper is a must-read. It draws a clear line in the sand: the gap between "can chat" and "can build" is currently a chasm.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →