AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution TracesExplained for Beginners
Sungho Park, Wonjoong Kim, Rongyuan Tan +10 more
Abstract
LLM agents remain unreliable on long-horizon tasks, where small local failures can compound over extended interactions and lead to overall task failure. Although external harnesses can substantially improve robustness, harness design remains a manual and expensive process that requires searching over a large space of prompts, tool configurations, and control logic. We propose AutoSaddler, an automatic harness optimization framework that formulates harness improvement as an offline learning problem and iteratively updates the harness using failure signals from mini-batches. AutoSaddler combines failure-trace diagnosis, structured patch generation that treats the harness as code, and validation-based update selection. Experiments on GAIA2, SWE-Bench Pro, and Terminal-Bench 2.0 show that AutoSaddler substantially improves agent performance over the corresponding base harnesses, achieving gains of 9.0, 9.6, and 10.0 percentage points, respectively. Ablation studies further suggest that effective harness optimization benefits from three ingredients: deep debugging rather than shallow reflection, targeted modifications rather than unconstrained editing, and generalization-aware selection rather than trajectory-specific repair. Together, these results suggest that automatic harness optimization is a promising path toward more performant and reliable agent systems.
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces
The Problem
LLM agents continue to struggle with long-horizon tasks—workflows that require many sequential steps, tool use, and environment interaction. In these settings, a single small mistake early in the process can compound over time, causing the entire task to fail. While "external harnesses"—additional layers of code, prompts, and logic wrapped around the LLM—can substantially improve robustness, designing these harnesses remains a manual, expensive process. Researchers must search through a massive space of prompts, tool configurations, and control logic, and evaluating each candidate is costly because the agent may need to execute many steps before success or failure becomes clear.
AutoSaddler addresses this gap by formulating harness improvement as an offline learning problem. Instead of manual tuning, it automatically iteratively refines the harness using failure signals from agent execution traces. The goal is to produce more performant and reliable agent systems without the need for exhaustive human search.
How It Works (The Technical Mechanics)
At a high level, AutoSaddler operates like a mini-batch machine learning loop, but for textual harness parameters. The framework runs in iterations. In each iteration, it evaluates the current harness on a mini-batch of tasks, diagnoses why tasks fail, generates structured patches to the harness, validates those patches, and then selects the best updates for the next round.
The Workflow
- Evaluation: The current harness is tested on a mini-batch of training tasks. Outcomes and execution traces are recorded.
- Diagnosis-Patch Session: Failed traces are fed to a specialized "Diagnosis-Patch Agent." This agent doesn't just do shallow reflection; it actively explores both the execution traces and the harness codebase to identify root causes. It produces a structured patch that targets specific components (prompts, tools, or middleware).
- Verification: The patched harness is re-evaluated on the same mini-batch. If performance improves, the patch is considered a "mini-batch improvement."
- Generalization Check: The patched harness is evaluated on a development set to see if the improvement generalizes beyond the specific mini-batch.
- Reflection & Evolution: Regardless of whether the patch was accepted, the system records lessons learned (what fixed, what regressed, etc.) into an EvoDAG (a directed acyclic graph that tracks the history of harness updates). An "Evolution Agent" then consults this graph to propose the next harness candidate , potentially merging successful components from previous iterations.
Key Analogies
Think of AutoSaddler like a software debugging loop for an application:
- The Harness is the Code: Just as a developer might fix a bug in a function, AutoSaddler patches the harness.
- The Mini-Batch is the Test Suite: Instead of running the whole program, it tests on a small batch of cases.
- Diagnosis is Bug Analysis: The agent looks at logs (traces) and the codebase to figure out why it crashed.
- The Patch is the Fix: A targeted change to the code.
- The Development Set is the Integration Test: Checking if the fix works across the whole system, not just the specific test case that failed.
- EvoDAG is the Git History: It remembers every change made, why it was made, and whether it helped or hurt, allowing the system to revert bad changes or cherry-pick good ones from different branches of history.
Structured Patching and Phased Scheduling
A critical design choice is how patches are generated. The harness is organized into three layers:
- Prompt: System instructions and rules.
- Tool: Available tools and their interfaces.
- Middleware: Runtime control logic (hooks, agent loops).
Rather than allowing the agent to edit the harness haphazardly, AutoSaddler uses a taxonomy that categorizes patches into Capability Patches (which modify executable code or orchestration logic, like adding a new tool or fixing a bug) and Steering Patches (which are textual edits that leave code unchanged, like modifying a prompt or a hook reminder).
The framework employs Phased Patch Scheduling: it starts by prioritizing Capability Patches to address fundamental gaps in tooling and infrastructure. Once the harness's functional capacity has stabilized, it transitions to Steering Patches to refine the agent's behavior within that established capability set. This is analogous to starting with large learning-rate steps and then moving to smaller ones as the model converges.
In-Depth Diagnosis
AutoSaddler distinguishes its diagnosis from "shallow reflection" by actively using the Claude Agent SDK to perform structured analysis. The agent reads the full execution trace (every tool call and response) and the agent codebase. It doesn't just ask an LLM "what went wrong?" based on a summary; it performs "history-analysis" and "diagnose" skills to pinpoint exact root causes. This deeper investigation is what allows it to find structural issues (like missing tools or incorrect logic) that a single LLM call overlooking a summary would miss.
Key Results & Benchmarks
AutoSaddler was evaluated on three major benchmarks: GAIA2, SWE-Bench Pro, and Terminal-Bench 2.0. In all cases, it substantially outperformed the base harness (the default agent setup without optimization).
Performance Gains
- GAIA2: Improved by 9.0 percentage points (from 53.0% to 62.0% Pass@1).
- SWE-Bench Pro: Improved by 9.6 percentage points (from 37.3% to 46.9% Pass@1).
- Terminal-Bench 2.0: Improved by 10.0 percentage points (from 40.0% to 50.0% Pass@1).
Furthermore, AutoSaddler exceeded the strongest automated baselines (GEPA and Meta-Harness) on each benchmark by 7.4, 4.4, and 4.4 percentage points respectively. On Terminal-Bench 2.0, it even surpassed a manually expert-tuned harness (Terminus KIRA) by 2.5 percentage points.
Efficiency Gains
A standout aspect of AutoSaddler is its efficiency. It achieves high performance with far fewer task-agent rollouts (the expensive "compute" of running the agent):
- GAIA2: AutoSaddler reaches 72.3% development accuracy after only ~1,000 rollouts, while other methods saturate around 64-61% despite using ~2,800 rollouts. When measured by the number of traces leveraged for learning, AutoSaddler reaches its best score after only 147 rollouts—roughly 10 times fewer than the strongest baseline (1,400 rollouts).
- Terminal-Bench 2.0: Starting from 52.6%, AutoSaddler reaches 73.7% dev accuracy after only 31 task executions and 12 leveraged traces, outperforming the next best method by over 10 percentage points.
Why It Matters (Key Takeaways)
- Automation of a Painful Process: AutoSaddler eliminates the need for manual harness tuning. By formulating harness optimization as an offline learning problem, it can adapt harnesses when switching to a new LLM or deploying in a new domain much more efficiently than human experts can.
- Durable, Generalizable Updates: The framework's three key ingredients—deep debugging, targeted modifications, and generalization-aware selection—work together to ensure updates aren't just "band-aids" for one specific task. The structured patching prevents the agent from overfitting to the specific failures of a mini-batch, and the EvoDAG-based evolution ensures that useful patterns are preserved across iterations.
- Significant Performance Improvements: Gains of 9-10 percentage points represent a substantial improvement in agent reliability. In practical terms, this means the agent succeeds on roughly 1 in 10 more tasks than before, which is significant for applications relying on autonomous task completion.
- Caution for Deployment: The paper notes that because AutoSaddler formulates harness optimization as a supervised learning problem, it assumes access to a training set of tasks with gold answers. In real-world deployment, such ground-truth labels may be scarce. The authors recommend that automated harness modifications should undergo human review or additional security validation before deployment to production systems.
Overall, AutoSaddler represents a promising path toward more robust LLM agents, demonstrating that automatic harness optimization can yield major gains in performance and efficiency.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →