Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task SelectionExplained for Beginners
Atsuyuki Miyai, Kiyoharu Aizawa, Toshihiko Yamasaki
Abstract
We present a novel approach to efficient LLM harness optimization through adaptive validation task selection. Harness optimization iteratively rewrites the harness code based on validation performance, enabling substantial performance gains without updating the underlying model weights. Existing approaches, however, evaluate a fixed validation set in full at every iteration, incurring substantial evaluation costs even on tasks that become less discriminative as the harness evolves. We propose Task-CoEvolve, which co-evolves the validation tasks with the harness by addressing two challenges: selecting informative tasks and estimating full-set performance from partial evaluations. Task-CoEvolve builds on the observation that tasks on which candidate harnesses disagree are more informative for distinguishing among them than tasks that are consistently solved or failed. It uses variance-weighted sampling based on past outcomes to focus evaluation on tasks near the capability frontier, with the sampling distribution adapting as the harness evolves. It then estimates full-set scores from the sampled tasks by accounting for their sampling probabilities, enabling consistent comparisons across iterations despite evaluating different subsets. Experiments on online text classification and Terminal-Bench 2.1 show that Task-CoEvolve consistently outperforms subset-based baselines and matches the final performance of full-set search while reducing the number of evaluations during optimization by 80%. Code will be released at https://github.com/Agent4Science-UTokyo/Task-CoEvolve.
Task-CoEvolve: Smart Evaluation, Bigger Gains
If you’ve ever waited for a language model to finish a complex task, you know that the "harness"—the code that frames how the model approaches a problem—can be just as important as the model itself. A clever harness can squeeze out six times more performance from the same underlying AI. But finding that clever harness manually is a nightmare of trial and error.
Enter Task-CoEvolve, a method that doesn’t just optimize the harness, but also optimizes how we evaluate the harness. The result? You get similar or even better AI performance while slashing evaluation costs by up to 80%.
Here is the breakdown of what this paper is doing, why it matters, and the key takeaways.
1. The Problem: The Cost of "Looking Everything Over"
Imagine you are trying to evolve a better harness for an AI. Every generation, you need to test the new harness design. In existing systems, the standard approach is to test every single candidate against the entire validation set of tasks.
This creates two major headaches:
- The Cost Explosion: If you have 89 complex tasks (like using an AI to book a flight or handle a customer support ticket), and you test 60 different harness designs over 20 iterations, you are running nearly 10,000 task executions. For expensive tasks—like long-horizon coding or terminal tasks that take minutes each—this becomes the bottleneck, dominating the entire optimization process.
- The Stagnation Trap: As the harness gets better, many tasks become "solved" by every candidate, or "unsolvable" by everyone. Testing these tasks continues to waste compute, even though they no longer provide useful information for distinguishing between good and bad harnesses.
The paper notes that simply resampling a fixed set of tasks every iteration doesn't fully solve this, because the raw scores from different subsets aren't comparable (a "lucky" easy subset vs. a "hard" subset).
2. How It Works: The Two-Step Strategy
Task-CoEvolve solves these problems with two clever mechanics working in tandem.
The Mechanics: Variance-Weighted Sampling
The core insight is simple but powerful: Tasks where candidates disagree are the most informative.
Think of it like a multiple-choice test. If every student gets a question right, the question doesn't help you rank the students. If every student gets it wrong, it also doesn't help. But if half the students get it right and half get it wrong, that question is gold—it distinguishes the strong from the weak.
Task-CoEvolve measures this using Bernoulli variance. For each task, it looks at the historical success rate of past harnesses. If a task has a 50/50 split of successes and failures, its variance is high, and it gets sampled heavily. If a task has a 99% success rate, its variance is near zero, and it gets ignored.
Critically, the "sampling distribution" adapts as the harness evolves. Early on, you might sample a broad set of tasks. As the harness improves, the distribution automatically shifts to focus only on the "capability frontier"—the tasks that are just barely solvable and thus provide the best signal for improvement.
The Math: Making Unfair Comparisons Fair
Here is the technical trick that makes the whole thing work. Because the method samples different tasks at every iteration, you can't just add up the scores directly—sampling a "hard" subset will always look worse than sampling an "easy" subset.
To fix this, Task-CoEvolve uses Horvitz-Thompson estimation (a statistics technique for sampling without replacement). It calculates the probability of each task being selected () and weight the results accordingly.
- The Hajek Estimator: Used when tasks are very easy or very hard (success rates near 0 or 1). It weights each outcome by .
- The Anchored Difference Estimator: Used for tasks in the middle. It compares the task outcome to a historical "anchor" rate, stabilizing the estimate.
In plain English: The method keeps a running tally of how likely each task was to be picked, and uses that to "correct" the score, ensuring that a candidate evaluated on a random subset of hard tasks is compared fairly against a candidate evaluated on a random subset of easy tasks.
3. Key Results: 80% Less Work, Same (or Better) Results
The experiments cover two domains: standard text classification and a tough benchmark of long-horizon tasks (Terminal-Bench 2.1).
Text Classification (The Quick Wins)
- 7% Budget: Task-CoEvolve achieved 47.6% accuracy. For context, a "Naive" baseline using 100% of the data hovered around 45.2%. Task-CoEvolve matched full-search performance using only 1/14th of the evaluation effort.
- 20% Budget: Task-CoEvolve hit 49.3%, actually beating the full-search baseline and the Meta-Harness baseline. It outperformed Meta-Harness by about 1%, suggesting that changing evaluation samples reduces overfitting.
Terminal-Bench 2.1 (The Hard Stuff)
This benchmark involves expensive, autonomous agent tasks.
- The Metric: Task-CoEvolve used only 20% of the tasks per iteration.
- The Cost Reduction: It reduced the total search cost by 67–80%.
- For one model (GPT-5.6 Luna), it dropped from 22.2 hours of compute to 11.5 hours.
- For another (Qwen3.6), it dropped from 38.0 hours to 20.5 hours.
- The Performance: The final performance was virtually identical to Full Search, losing only about 1 percentage point (roughly one task out of 89).
The "Gotcha" on Cost
The paper points out an interesting nuance: Not all 20% subsets are created equal.
- Random sampling of 20% tasks might be cheap because the tasks are easy and finish quickly.
- Task-CoEvolve's smart sampling focuses on the "hard" tasks where candidates disagree.
- Consequently, Task-CoEvolve actually uses more tokens per trial (3.2M) than Random sampling (0.7M). However, because it is so much more efficient at finding good harnesses, the total search cost is still 5x lower than Full Search.
4. Why It Matters: Key Takeaways
Here are the four most important things to walk away with:
- Evaluation is the bottleneck: In harness optimization, how you evaluate is often more limiting than how you generate candidates. This method proves you can drastically cut that cost.
- Adapt or stagnate: The method’s ability to shift its focus as the harness improves is key. It prevents the "solved tasks" trap, ensuring that the evaluation signal stays sharp throughout the optimization process.
- Fair comparison is possible: By using the Horvitz-Thompson correction, the method allows you to compare candidates evaluated on different subsets. This means you can explore the harness space more broadly without breaking your ability to select the winner.
- The "Early Stop" Limitation: The paper notes a current limitation: the method fixes the number of tasks to evaluate per candidate upfront. It cannot "give up" early on a clearly bad candidate, nor can it spend extra time deliberating on ambiguous ones. Future work will likely focus on making this adaptive.
The Bottom Line
Task-CoEvolve offers a practical path to more efficient AI development. By treating the validation set as something that should evolve alongside the harness—and by using statistical tricks to make partial evaluations comparable—it slashes the compute cost of optimization by up to 80% while maintaining or improving the final quality of the system. For product managers and engineers, this means faster iteration cycles and lower compute bills for the same (or better) AI performance.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →