arXiv:2609.08183nvidia/nemotron-3.5-lightning-30b-a3bSeptember 8, 2026

NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing HarnessExplained for Beginners

NeoHorse Team, Guoliang Cao, Guohao Dai +34 more

Computation and Language

Abstract

Recursive self-improvement (RSI) requires a concrete mechanism through which an AI system observes its capabilities and converts that evidence into the next round of learning. We present NeoHorse-1, a family of agent-native models developed to explore this path through agentic post-training. Our system combines a heterogeneous model pool with intelligent routing, recording the predicted capability demand, selected service tier, and subsequent interaction for each user turn. These records are converted into training examples that preserve interleaved reasoning, tool calls, and harness context, and are admitted through structural validation, six-dimensional semantic evaluation, and subscene-level labeling. Routing signals organize supervised fine-tuning into a three-stage curriculum and extend to routing-guided on-policy distillation, where a teacher supervises student-generated responses under the same progression. Capability-guided allocation then converts evaluation feedback into the next training mixture, closing an evaluation-selection-update loop in which what the system learns to do shapes what it learns from next. Across eleven benchmarks covering harness-based agents, tool use, coding, and instruction following, post-training raises the macro-average from 58.94 to 64.87 at 4B and from 65.60 to 69.04 at 9B, substantially narrowing the aggregate gap between the post-trained 4B model and the 9B base model. NeoHorse-1 provides an initial prototype of this feedback-driven process and a path toward harness-mediated RSI across successive iterations.

The Problem

Most of us interact with AI through static prompts: we type a question, and the model does its best to answer. But building a system that gets smarter over time requires a different trick. In the research world, this is called “recursive self-improvement” (RSI)—the idea that an AI can look at its own performance, notice where it’s weak, and then use that self-diagnosis to fuel its next training run.

The gap is practical. AI labs have made huge strides in pre-training (the massive upfront ingestion of text and code), but once a model is deployed, guiding its ongoing evolution is surprisingly manual. Researchers often rely on human annotators to write new examples, or they use crude reward models that can’t capture the richness of real-world tasks. There hasn’t been a clean, automated loop where the model essentially says, “Hey, I struggled with this type of problem; give me more like it.”

NeoHorse-1 attacks this head-on. The authors argue that to achieve true RSI, we need a system that can observe its own behavior in the wild, package those observations into training data, and then re-inject them into the model in a structured way. They focus on “agentic post-training”—teaching a model not just to answer questions, but to use tools, reason through steps, and decide when to call external services. The goal is a feedback loop: the model’s actions become the curriculum for its next version.

For someone outside academia, the problem is this: we want AI assistants that don’t just know a lot, but that can adapt, learn from their mistakes, and get better at the specific things we ask them to do—without us having to constantly hand-craft new training data. NeoHorse-1 is a prototype attempt to automate that cycle.

How It Works (The Technical Mechanics)

Let’s dive into the machinery without getting lost in the weeds. The paper describes NeoHorse-1 as a family of models that rely on three core ideas: a heterogeneous model pool, intelligent routing, and a curriculum of supervised fine-tuning.

Imagine you have a small fleet of models: some are “smart but slow” (large, compute-heavy), and some are “fast but lighter.” When a user sends a query, an intelligent router decides which model should handle it. The router isn’t just picking randomly; it’s predicting the capability demand. Is this a simple factual question? Send it to the lightweight model. Is it a multi-step coding task? Route it to the heavyweight.

Here’s the clever part: every single routing decision, every tool call, every intermediate reasoning step is recorded. Think of it as a black-box flight recorder for AI interactions. These records aren’t just discarded; they’re converted into training examples. The paper emphasizes “interleaved reasoning, tool calls, and harness context.” In plain English, if the model reached for a calculator, that action and the result are saved alongside the final answer.

But raw recordings aren’t enough. The paper introduces structural validation and a six-dimensional semantic evaluation. I like to think of this as a quality filter. Before a recorded interaction is used to train the next model, it gets scored across dimensions like correctness, coherence, and tool usage quality. Only the “golden” examples pass muster.

Now, how does the actual learning happen? The authors use a three-stage curriculum for supervised fine-tuning (SFT). It’s analogous to a student progressing through a course:

  1. Foundation: The model learns basic instruction following and simple tool use.
  2. Intermediate: It tackles more complex reasoning, perhaps requiring multiple tool calls.
  3. Advanced: It handles the most demanding harness-based agent tasks.

This curriculum is crucial. You wouldn’t start a beginner driver on a Formula 1 track; similarly, training a model on ultra-hard tasks from day one usually yields poor results.

The paper also introduces “routing-guided on-policy distillation.” This is a fancy term for a teacher-student setup where the teacher is the current iteration of the model (or a stronger baseline), and the student is the one being trained. The twist? The student is supervised to generate responses under the same routing progression. If the router sends a hard problem to the teacher, the student is expected to mimic that behavior, effectively learning the “art” of routing and tool use.

Finally, there’s “capability-guided allocation.” This closes the loop. After the student is fine-tuned, it’s evaluated again. The feedback from that evaluation—what it got right, what it got wrong—feeds back into deciding what the next training mixture should look like. It’s a cycle of evaluate-select-update.

Key Results & Benchmarks

The numbers are the part that usually grabs attention. The paper reports results across eleven benchmarks covering harness-based agents, tool use, coding, and instruction following. The key metric is a macro-average score.

Here is the before-and-after at two model sizes:

  • 4B parameter model: The macro-average score jumps from 58.94 to 64.87.
  • 9B parameter model: The score climbs from 65.60 to 69.04.

What does this mean in plain language? The post-trained 4B model narrows the performance gap with the 9B base model. Originally, the 9B model was significantly ahead. After this post-training process, the 4B model is much closer—the gap is substantially narrowed. It’s like taking a smaller, faster car and, through a round of targeted upgrades, making it competitive with a larger, heavier model in its class.

Across the board, the improvements are consistent. The paper doesn’t just highlight one magic benchmark; the lift happens across the board—harness agents, tool use, coding, and following instructions. This suggests the recursive loop is working: the model is getting better at the things that matter for general capability, not just memorizing answers to specific test questions.

Why It Matters

The implications are wide-ranging, and the paper leaves us with a few concrete takeaways:

  • Automated Alignment and Adaptation: If this feedback loop can be scaled, it means AI systems could potentially adapt to new domains without waiting for a new pre-training run or a team of human labelers. Imagine an AI assistant that gets better at coding Python the more you use it, simply because it logs your interactions and retrains itself overnight.
  • The Data Flywheel: The paper effectively proposes a “data flywheel.” Real user interactions generate data → the model improves → better interactions generate better data. This is the holy grail for building ever-more-capable systems without hitting a data bottleneck.
  • Routing as a Feature, Not a Bug: A fascinating insight is that the routing mechanism itself becomes a training signal. By observing which model handles which task, the system learns what “hard” looks like. This could lead to more efficient compute usage—only spending heavy compute when truly necessary.
  • Limitations to Watch: As with any RSI prototype, there’s a risk of “reward hacking.” If the evaluation metrics aren’t perfectly aligned with human intent, the model might optimize for the wrong things. Also, the paper notes this is an initial prototype. We’re far from a fully autonomous, recursively improving AI, but this paper charts a promising path toward it.

What to watch for next: Likely, follow-up work will focus on scaling this to much larger models, improving the stability of the feedback loop, and exploring how these techniques interact with Reinforcement Learning from Human Feedback (RLHF), which is currently the standard way AI assistants are tuned to be helpful and harmless.

NeoHorse-1 isn’t the final word on recursive self-improvement, but it provides a concrete, working prototype of how we might close the loop between a model’s actions and its continued learning. It’s a significant step toward AI that doesn’t just sit there knowing things, but actively gets better at using them.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →