arXiv:2608.24876nvidia/nemotron-3.5-lightning-30b-a3bAugust 25, 2026

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent HarnessesExplained for Beginners

Zhaochen Yu, Yingcheng Wu, Zhenfei Yin +5 more

Artificial IntelligenceComputation and Language

Abstract

Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation. We introduce Recuris, a recursive Experiential-Working Memory architecture for long-horizon agent harnesses, in which Working Memory tracks task progress and guides skill selection from Experiential Memory, grounding skill use in current needs rather than the full history. This coupling also turns execution into structured evidence that localizes failures to specific memory components. Across tasks, a fixed Meta-Agent turns that evidence into localized, validation-gated updates to Skill Memory that reshape execution and yield new evidence, forming a bounded recursive memory-evolution loop. Across four long-horizon benchmarks and ten models, Recuris improves task success in 35 of the 37 completed model-benchmark pairs, carrying frontier models to SOTA-level task success: on tau-bench it adds +17.8 points to GPT-5.6 Sol and +15.6 to Claude Opus 5, taking Opus 5 to 87.9%, and +16.6/+13.5 points on Qwen3.6-27B/35B on SkillFlow. The advantage widens as the interaction horizon grows, to +32.2 points on the longest tasks, and common long-horizon failures fall by up to 80%. These results position recursively evolving memory as a scalable foundation for RSI, enabling agents to continuously transform accumulated experience into increasingly effective long-horizon behavior. Code: https://github.com/Gen-Verse/Recuris

The Problem

Recursive self-improvement (RSI) is notoriously difficult for long-horizon AI agents. The core issue is that as tasks stretch over many steps, the agent’s growing interaction history becomes a blur of completed steps, outdated information, and execution noise. Traditional harnesses rely on retrieving skills from either the initial prompt or the full chat history, but both approaches fail in practice. The initial instruction quickly becomes misaligned with the current problem, while searching the entire history forces the model to sift through irrelevant completed actions and noise. The fundamental gap isn't a lack of useful experience—it's the absence of a compact, reliable "task state" that can continuously align stored experience with the agent's real-time needs. Without such a state, agents lose track of unresolved goals and invoke obsolete or misaligned skills, causing failures to compound across long executions.

How It Works (The Technical Mechanics)

Recuris introduces a recursive Experiential–Working Memory (EM–WM) coupling that solves this alignment problem. The architecture hinges on two interacting memory systems:

Working Memory (WM) acts as a compact, verified task-state tracker. At every step, it records the agent's current progress: which goals are pending, done, or blocked, and what observations support that progress. Crucially, WM only records changes that are verified by actual environment feedback, not the model's own claims. This creates a trustworthy "state of what remains to be done."

Experiential Memory (EM) stores reusable skills—modular pieces of knowledge or behavior the agent has encountered before.

The innovation is how these two interact. Instead of retrieving skills from the full history, the harness uses WM's current state to decide when and which skill to invoke from EM at defined execution events (e.g., when a tool call is drafted, or at turn boundaries). After a skill executes, WM updates its state based on the actual environment result, and only those changes supported by evidence are committed.

This creates a closed loop: Current Task State → Skill Selection → Execution Feedback → Updated Task State. Because skill invocation is grounded in the verified current state rather than the full history, the agent stays focused on what actually needs doing right now. This also produces a valuable side effect: execution becomes structured diagnostic evidence. Every step records which task state triggered which skill, what action was taken, and what observation resulted. When a task fails, this trace tells the system exactly which memory component—WM, EM, the invocation policy, or the verification checker—bears responsibility for the failure.

The recursion happens across tasks. A fixed "Meta-Agent" analyzes these structured traces from failed runs, diagnoses the failing component, and proposes a localized patch to that component only. A fixed validation gate then tests the patch on a held-out set of anchor tasks: it must repair the target failure without regressing on tasks the memory already solves. If the patch passes, it updates the Skill Memory; if not, the memory stays unchanged. This creates a bounded recursive loop: memory shapes behavior, behavior produces evidence, evidence evolves the memory—all without ever touching the underlying model weights or the full agent program.

A concrete analogy: Imagine a software team debugging a complex application. Without WM, a developer might search the entire git history for clues every time a bug appears—the codebase grows, and the search becomes inefficient and error-prone. With WM, the developer maintains a concise "current state of what's broken" ticket. When a new bug emerges, they check that ticket first, look only at the relevant recent changes (EM), and fix just the component indicated. The "structured trace" is the git commit history that tells them exactly which line of code caused the issue. Recuris automates this debugging process for AI agents, but across tasks rather than within a single codebase.

Key Results & Benchmarks

The empirical results are striking and directly address the long-horizon problem. Across four benchmarks and ten models (spanning 3-billion-parameter open-weight models to frontier systems), Recuris improves task success in 35 of 37 model-benchmark pairs. The gains are substantial and counter-intuitive: they increase with task length rather than decaying.

On the τ2-Retail benchmark, Recuris lifts GPT-5.6 Sol from 58.3% to 79.1% (+20.8 points) and Claude Opus 5 from 72.4% to 87.9% (+15.6 points). On SkillFlow, it raises Qwen3.6-27B from 42.2% to 58.7 (+16.5 points) and Qwen3.6-35B from 35.3% to 48.8 (+13.5 points). Critically, the advantage widens on the longest tasks—reaching +32.2 points on the most horizon-intensive evaluations—and common long-horizon failures (like omitted writes and error cascades) fall by up to 80%.

Perhaps most importantly, these gains transfer across models. A single Evolved Skill Memory, distilled from a mid-sized deployment model, lifts every model it's applied to, including frontier models it never saw during evolution. On τ2-Retail, the same package adds +17.8 points to GPT-5.6 Sol and +15.6 to Claude Opus 5. On SkillFlow, it transfers broadly with scale, yielding +16.6 and +13.5 points on the two largest Qwen3 models. The memory carries "discipline"—knowledge of what to verify and when a goal is still open—rather than raw capability, which is why it benefits stronger models that already have ample capability but struggle with long-horizon coordination.

The paper also quantifies why Recuris works. Structured traces enable component-level failure localization at 64.8% accuracy versus only 13.0% when judging from task outcomes alone. This precision is what makes scoped repairs possible. The harness on its own contributes nothing measurable; all gain comes from what the evolution loop admits. And crucially, because recursion is confined to the external memory-control layer, the base model remains exactly as the provider shipped it—no fine-tuning, no weight updates.

Why It Matters (Key Takeaways)

  • Recursive self-improvement becomes feasible. By confining evolution to an external memory-control layer, Recuris enables continuous improvement of long-horizon agent behavior without ever modifying the underlying model. This is a scalable foundation for RSI: agents can transform accumulated experience into increasingly effective behavior over time.

  • Memory must be state-grounded. The results demonstrate conclusively that simply adding more context or skills to a prompt does not help—and can actually hurt (the model-controlled variant, which floods the prompt with all skills, scores 18 points worse than Recuris). What matters is whether the agent knows when a skill is needed, and that determination comes from a verified working state.

  • Failures localize precisely. The structured EM–WM coupling turns execution into diagnostic evidence, allowing the system to attribute failures to specific memory components (WM, EM, invocation policy, or checker) with 64.8% accuracy versus 13.0% from outcome alone. This precision enables targeted repairs rather than wholesale rewrites.

  • What to watch next. The paper notes that transfer is narrow: a memory evolved from failures on one type of task transfers best to tasks still containing that kind of failure. The τ2-Airline lineages show no transfer to held-out tasks because the evolution runs already repaired all improvable failures. Additionally, the per-attempt success gains, while directional, haven't yet reached statistical significance at current sample sizes—but the within-task effects (e.g., solving a task 15/16 times with adapted memory vs. 7/16 with seed memory) suggest substantial reliability improvements on individual hard tasks.

Recuris positions recursively evolving memory as a practical mechanism for long-horizon agent improvement. It doesn't require smarter models; it requires better memory orchestration. By grounding skill invocation in verified task state and enabling component-level evolution through structured evidence, it opens a path toward agents that continuously get better at long tasks the more they run—exactly the recursive capability the field has been seeking.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →