ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style HoldoutsExplained for Beginners
YuanHang Xiao
Abstract
Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue this under-specifies what is being evaluated: the proper unit is a declared model-plus-runtime configuration whose failures can occur in evidence acquisition, runtime routing, safety boundaries, or repeated execution. We present ClawProBench, a trace-aware benchmark for runtime-native agent evaluation instantiated on OpenClaw, a live agent runtime with workspace tools and native surfaces for browsing, memory, messaging, scheduling, skills, and subagents. ClawProBench defines two tracks: a 102-scenario full profile with live workspace and native-runtime routing tasks, and a frozen 68-scenario holdout with closed-world JSON output contracts for robust ranking. Trials are scored from execution traces via a safety-gated formula combining correctness, process quality, and efficiency, preserving failure evidence for audit. Our anonymous artifact includes benchmark definitions, scoring code, manifests and sanitized traces. We evaluate 68 configurations on the full profile and 37 on holdout. The top safety-gated average trace score is 0.7671. Native-runtime tasks underperform workspace-live tasks (0.5238 vs. 0.6415). On holdout, pass@k-any outperforms strict three-trial pass (0.6638 vs. 0.2890), while full-profile and holdout rankings show weak alignment (Spearman 0.1300). Rankings based purely on correctness differ substantially from process-aware, safety-gated and strict-pass views. Final-answer leaderboards may hide native-surface weaknesses, one-off successes and trace-local agent failure modes.
The Problem: Evaluating the Loop, Not Just the Answer
If you’ve ever watched an AI agent struggle to find a file it just created, or loop endlessly on a browser tab, you’ve felt the gap this paper identifies. Most benchmarks today are "final answer" machines. You give the agent a task, and the benchmark simply checks: Did it get the right result at the end?
But as the authors argue, this under-specifies what’s actually happening. When an agent operates in a stateful runtime—reading files, browsing the web, sending messages, scheduling reminders—a correct final answer can mask a chaotic internal process. The model might have gotten lucky, used the wrong tool, skipped critical evidence, violated a safety boundary, or left no audit trail for a user to inspect later.
The real unit of evaluation, the paper contends, is the declared model-plus-runtime configuration. Failures can happen at multiple points: evidence acquisition (did it find the right data?), runtime routing (did it call the right surface?), safety boundaries (did it overstep?), or repeated execution (can it reliably do it again?). Existing benchmarks collapse all of this into a single success/failure boolean. ClawProBench steps in to measure the entire pipeline.
How It Works: The "Claw" Runtime and Trace-Aware Scoring
To make this measurable, the authors built OpenClaw, a live agent runtime with native surfaces for browsing, memory, messaging, scheduling, skills, and sub-agents. They then designed ClawProBench, a benchmark that evaluates a "configuration" (model + runtime + tools + safety filters) rather than just the model itself.
The benchmark has two tracks:
- The Full Live Profile (102 scenarios): A demanding set of tasks covering workspace work (66 tasks) and native OpenClaw surfaces (36 tasks) like skills, browser, memory, messages, sessions, directory, cron, and agent delegation.
- The Frozen Holdout (68 scenarios): A "closed-world" set of tasks with fixed JSON output contracts. This freezes the scenario identities and checking logic, allowing researchers to run the same tasks across different runtimes (OpenClaw, IronClaw, NanoClaw) for fair comparison.
The Scoring Formula: Looking at the Trace
The core innovation is how trials are scored. Instead of a binary pass/fail, each execution trace is graded using a safety-gated formula:
Here is the breakdown, translated into plain language:
- (Correctness): Did the agent get the end state and artifacts right? (0 to 1)
- (Process Quality): Did it use the required tools in the right order? Did it avoid redundant steps? This rewards "good behavior" even if the final answer is slightly off.
- (Efficiency): A penalty for excessive tool calls. If the optimal number of steps is 5 and the agent took 10, it gets penalized.
- (Safety Gate): A non-compensatory check. If the agent violates a safety boundary (e.g., accessing private data), the score drops to zero, regardless of how correct the answer is.
Analogy for the PM/Engineer: Think of this like a CI/CD pipeline deployment. A traditional benchmark only checks: "Was the website live at the end?" ClawProBench, instead, checks the deployment log. It asks: "Did the deployment use the correct server surface? Did it skip the health check step (process penalty)? Did it take 10 minutes when 2 were needed (efficiency penalty)? And crucially: Did it accidentally delete the production database? If so, the safety gate zeros out the entire score, no matter how pretty the website looks."
Key Results: Native Struggles, Reliability is Hard, and Leaderboards Lie
The empirical findings paint a nuanced picture that challenges how we currently view AI agent leaderboards.
- Native vs. Workspace: On the full profile, native-runtime tasks (browser, skills, etc.) scored significantly lower (0.5238) than workspace-live tasks (0.6415). This suggests that current models are much better at manipulating local files and environment than they are at navigating native runtime surfaces like browsing or messaging.
- The Reliability Gap (The Holdout): On the frozen holdout, there is a stark difference between solving a task once and solving it reliably. The "pass@k-any" (success in at least one of three trials) was 0.6638, but the strict "three-trial pass" (success all three times) was only 0.2890. This means many models can "get lucky" once, but cannot replicate the success consistently—a critical insight for product reliability.
- Rankings are Unstable: The paper found weak correlation between the full profile rankings and the holdout rankings (Spearman correlation of 0.1300). A model that tops the full 102-scenario profile might not perform well on the 68 specific frozen tasks, and vice versa. Furthermore, rankings based purely on final correctness differ substantially from those that consider process quality and safety.
- The "Hidden" Weaknesses: The authors demonstrate that final-answer leaderboards can hide critical weaknesses. A model might look competent because it gets the right answer, but it might do so by bypassing a required native surface, or it might only succeed on one specific run before crashing on the next.
Why It Matters: Key Takeaways
This research matters because it shifts the goalpost for AI evaluation. Here are the four key takeaways:
- Runtime Surface is a First-Class Metric: We can no longer treat "the model" as a monolith. Performance varies dramatically between workspace tasks and native surfaces (browser, memory, etc.). Evaluations must report surface-specific scores.
- Reliability ≠ Single Success: A high pass@1 score is not enough. For any practical application, we need to know if the agent can succeed consistently (pass@3 or strict pass). A model that solves a task 1 out of 3 times is far less useful than one that solves it reliably.
- The Safety Gate is Critical: The "safety-gated" scoring approach ensures that an agent cannot trade safety for performance. In a real-world product, an agent that "hacks" its way to the right answer by bypassing safety boundaries is a liability, not an asset.
- Trace Evidence is the New Gold Standard: By preserving execution traces and status labels, ClawProBench enables auditability. We can see why a model failed—was it a routing error? A missing piece of evidence? This moves the field from "did it work?" to "how and why did it work (or not)?"
Final Thought
ClawProBench represents a necessary maturation of AI benchmarking. It refuses to let a single correct answer obscure the complexity of the runtime environment. By demanding we look at the trace, the surface, and the reliability, it pushes the field toward building agents that are not just smart, but robust, safe, and consistent.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →