arXiv:2608.23200nvidia/nemotron-3.5-lightning-30b-a3bAugust 24, 2026

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow TasksExplained for Beginners

Xiao Zhang, Qumeng Sun, Jihao Li +4 more

Computation and Language

Abstract

Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end verification. Yet successful execution experience is typically lost after a single run, forcing subsequent models to rediscover strategies and failure modes from scratch. We study whether such experience can instead be externalized and reused through EvoMap, where verifier-confirmed execution trajectories are consolidated into structured Gene. To evaluate this setting, we introduce the Long-Workflow Benchmark (LongWoF-Bench), comprising 778 machine-verifiable tasks across code generation, agent-environment synthesis, mathematical reasoning, and rule following. On the 252 tasks with verifier-confirmed Opus trajectories, evolved EvoMap Gene outperform Skill across all seven evaluated models by 8.7-15.5 percentage points, with the gains extending to consumer models from different model families. In contrast, reference-distilled Gene do not exhibit the same advantage, indicating that compact representation alone is insufficient and that Gene utility is closely associated with verified experience provenance. For Claude Opus, Gene reuse also completes 39 more tasks than Skill while reducing solve-time token consumption by 9.9%. Together, these results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.

LongWoF-Bench: When Verified Experience Becomes Reusable Intelligence

1. The Problem: The Amnesiac Model Problem

Imagine you’re tackling a complex, multi-step project at work—a software launch, a data pipeline, or a research experiment. You muddle through, make some decisions, hit a few dead ends, eventually get it verified and working, and then close the ticket. The success is documented, perhaps in a wiki, but the specific sequence of choices that kept the whole thing from collapsing? The exact way you handled that one edge case boundary condition? That’s usually lost.

In the world of Large Language Models (LLMs), this is the "amnesiac model problem." Today’s most capable models are expected to execute complex workflows: generating functioning code, synthesizing agent-environment systems, solving multi-step math problems, or following intricate rules. Success in these "verifiable long-workflow tasks" isn't about whether individual steps look reasonable. It’s about whether the entire execution satisfies strict, interdependent constraints end-to-end. A single misplaced bracket, an incorrect default parameter, or an ordering mistake early in the process can invalidate the entire output, even if the later parts look perfect.

However, when a model successfully completes such a task, that hard-won execution experience is typically discarded. The next time the same—or a similar—task comes around, the model must rediscover the successful strategies and, crucially, the failure modes from scratch. It’s as if a master carpenter built a flawless cabinet, then burned the blueprints and forgot how they joined the wood, forcing every future carpenter to start with raw timber and guesswork.

Why should you care outside academia? Because this is an efficiency and reliability bottleneck. In production systems, AI agents are expected to handle complex, repetitive, or variant tasks. If every new attempt requires the model to "figure out" the workflow anew, it costs compute time, increases the risk of errors, and makes it hard to build on prior successes. The paper argues that we need a way to externalize and reuse this execution experience, turning a transient outcome into a reusable resource.

2. How It Works: EvoMap Genes vs. Skills

The paper introduces a framework called EvoMap to solve the amnesiac problem. The core idea is to treat successful model execution as a source of reusable knowledge, similar to how a library stores books.

The Two Types of Reusable Knowledge

The authors distinguish between two kinds of reusable guidance:

  • Skill: Think of a Skill as an IKEA-style instruction manual or a recipe. It codifies procedural knowledge: "First do X, then do Y, use tool Z." Skills are great for general procedures but abstract. They tell you how to do things in a general sense, but they might miss the specific, gritty details that actually make the difference between success and failure in a particular task.
  • EvoMap Gene: This is the star of the paper. A Gene is distilled from a verifier-confirmed execution trajectory. If a model successfully completes a task, the system records the entire path it took, including the corrections it made when things went wrong. This "execution-critical knowledge"—strategies, prerequisite checks, boundary conditions, and failure guards—is then packaged into a Gene.

The Analogy: If a Skill is a general recipe ("Boil water, add pasta"), a Gene is the specific chef's notes from a successful run: "Boil water at altitude X for Y minutes because the pasta brand Z absorbs water differently, and don't add salt until the water bubbles, or the sauce will break." The Gene preserves the why and the specific how that verified the result.

The Construction Process (Evolver)

Genes aren't manually written; they are evolved using a framework called Evolver. Here is the simplified process:

  1. Attempt: A model tries to solve a task.
  2. Feedback: If it fails, it receives sanitized feedback from a verifier (a automated checker that knows the correct answer or constraints).
  3. Refine: The model revises its solution.
  4. Iterate: This loop continues within a budget of attempts until a "passing trajectory" is achieved.
  5. Distillation: Once verified, the critical execution details from that successful path are extracted and packaged into a Gene.

Crucially, the paper notes that the Gene preserves the strategies and corrections that actually worked under verification, not just generic instructions.

Reuse via EvoMap

Once a Gene is created, it can be shared through the EvoMap infrastructure. A consumer model—potentially a completely different model family—receives the public task specification along with the Gene. It doesn't replay the original model's steps, but it uses the Gene as a guide. This separates the "cost of discovery" (the first hard run) from the "cost of application" (subsequent runs).

3. Key Results & Benchmarks: The Provenance Advantage

The LongWoF-Bench study put this to the test across 778 tasks spanning code generation, agent synthesis, math, and rule following. Here are the headline numbers.

The Main Event: Evolved Gene vs. Skill

On the 252 tasks where the producer model (Claude Opus 4.8) successfully evolved a verifier-confirmed Gene, the results were striking. Evolved Gene consistently outperformed Skill across all seven evaluated consumer models.

The gains were significant and consistent: 8.7 to 15.5 percentage points improvement in strict pass rate. To put this in perspective, if a model was passing 50% of tasks with just a Skill, adding the Gene bumped it up to roughly 64-66%.

This advantage extended beyond the producer model. Gains were observed across model families including Sonnet, Gemini, MiniMax, and Qwen. This suggests that verified execution experience is a form of "universal" knowledge that transfers between different AI systems, much like a mechanic's manual works across different car models.

The Critical Distinction: Provenance Matters

Perhaps the most important finding concerns reference-distilled Genes. On the 526 tasks where Opus failed to evolve a verifier-confirmed trajectory and instead used a "fallback" Gene distilled from teacher signals, the results flipped. Reference-distilled Genes trailed Skill for every model by 3.3 to 11.3 points.

This is a crucial insight. It shows that simply compressing knowledge into a compact Gene format is not enough. Without the direct, verifier-confirmed execution trail, the Gene lacks the evidence of which choices mattered. The paper summarizes this well: "compact representation alone is insufficient and that Gene utility is closely associated with verified experience provenance."

The Common Ground: Opus vs. Gemini

When comparing Genes produced by different models on the same 180 tasks, Opus-authored Genes consistently outperformed Gemini-authored Genes for every consumer model, with gains of 4.4 to 11.7 percentage points. This reinforces that the quality of the originating execution experience matters—better exploration and selection during the evolution process lead to more useful Genes.

The Efficiency Gain: Claude Opus Case Study

For the specific case of Claude Opus reusing its own Genes, the benefits were both quantitative and qualitative:

  • More tasks completed: Gene reuse completed 39 more tasks than Skill.
  • Token efficiency: It reduced solve-time token consumption by 9.9%. This means the model achieved more with less "thinking" cost, effectively amortizing the discovery cost across multiple runs.

4. Why It Matters: Key Takeaways

This research shifts how we think about LLM capabilities and deployment. Here are the four key takeaways:

  1. Verified Experience is a Differentiator: Simply having a procedure (Skill) is not as effective as having a verified execution history (Gene). The "provenance"—the fact that the knowledge was confirmed by a verifier—is what makes Gene valuable. This suggests that in real-world deployments, tracking how a task was solved successfully is as important as the solution itself.
  2. Cross-Model Transfer is Possible: Genes can transfer between different model families. An Opus Gene can help a Gemini or Qwen model. This opens the door to shared knowledge bases where the "hard-won experience" of one system becomes a productivity boost for others, potentially reducing the overall compute cost for an organization.
  3. The Cost of Discovery vs. Reuse: The study quantifies the efficiency gain. Multi-round discovery to find verified trajectories is expensive (many model calls and tokens). One-shot Gene reuse is significantly cheaper. This means that as Gene libraries grow, subsequent task execution becomes dramatically more efficient, addressing the concern that AI deployment becomes exponentially more expensive over time.
  4. Task-Specific Utility: The benefit of Genes varies by task type. They shine brightest in "agent-environment synthesis" and "rule following" tasks—areas where maintaining constraints, order, and boundary conditions is fragile and critical. In "code generation," the advantage is more model-dependent. In "mathematical reasoning," Genes help transfer strategies, but they can't fully compensate if the underlying model's reasoning capability is the bottleneck.

Summary

LongWoF-Bench provides evidence that we don't have to accept the amnesiac nature of current LLM workflows. By evolving and sharing EvoMap Genes—structured packages of verifier-confirmed execution experience—we can improve long-workflow completion rates, transfer knowledge across different AI systems, and reduce the computational cost of repeated problem-solving. The key takeaway is that how a model succeeds matters as much as that it succeeds, and that successful execution experience is a reusable resource worth capturing.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →