arXiv:2609.04304nvidia/nemotron-3.5-lightning-30b-a3bSeptember 3, 2026

Iris: Climbing to the Search FrontierExplained for Beginners

Ziyuan Liu, Hengqi Liu, Zichuan Wang +6 more

Artificial Intelligence

Abstract

We present Iris-mini and Iris-pro, two search agents trained at the 35B-A3B and 397B-A17B scales, together with the data pipeline and training recipe behind them. Tasks are reverse-constructed from the hyperlink structure of a web corpus: we author multi-hop chains over an entity graph distilled from a seed page and its out-links, rewrite every non-answer entity into a descriptive reference so that no clue can be resolved by string matching, and admit only questions that a reference model fails closed-book yet solves once the supporting evidence is supplied. These questions are then turned into trajectories, which are filtered at both the trajectory and the turn level before SFT. The policy is then optimized by RL against live search, with the reward judge and the observation summarizer served inside the training cluster, and with over-long rollouts interrupted at the request level and resumed from their committed prefix at the next step. We alternate the two stages in a procedure we call SFT-RL climbing, returning the hardest solved and most efficient rollouts of each RL round to the next supervised pass. Because inference-time context management is worth more on these benchmarks than most reported differences between systems, we evaluate every benchmark both with and without it, holding the tool set, the context limit, and the judge fixed. All results come from a single ReAct agent, with no sub-agents and no test-time verification. With management enabled, on BrowseComp, BrowseComp-ZH, DeepSearchQA, and HLE the two models reach 82.2/84.8/86.9/52.3 and 88.6/85.1/92.9/56.4, the strongest overall results among open-source search agents in their respective parameter ranges. We plan to release the model weights together with the complete recipe for data construction, training, and evaluation.

The Problem: Why Building a "Search Agent" Is Harder Than It Looks

Imagine you are trying to answer a question like, "Which city hosts the oldest surviving annual marathon that was founded after the fall of the Berlin Wall?" To a human, this is a moderate research task. To a large language model (LLM) without tools, it might seem impossible—the model only knows what was in its training data cutoff. Even with internet access, the model faces a fundamental problem: it must decide what to search for, how to interpret the results, and when to stop searching and give an answer.

Most current AI models are evaluated on "closed-book" reasoning, where the task, context, and computation budget are fixed in advance. Search agents break this mold by interacting with external tools. However, building a reliable search agent is surprisingly difficult. Recent research has focused on giving models tools, but substantially different performance can still arise from differences in the "inference-time harness"—the surrounding software that manages context and evaluation—rather than from the model's actual reasoning ability.

The gap this paper addresses: Prior work often reported results only with context management (CM) enabled. Since long search sessions can exhaust the available context window before the agent finds the answer, CM acts like a safety net that resets the conversation history. While helpful, this makes it hard to tell: Is the model smart, or just has a good editor? This paper aims to isolate the model's intrinsic search capability by evaluating performance both with and without CM, holding the tools and context budget fixed.

How It Works: The "SFT–RL Climbing" Engine

The authors propose a training pipeline they call SFT–RL climbing. It is an iterative loop combining two classic AI techniques: Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL).

1. Building the "Hard" Questions (Data Construction) The paper doesn't just use random questions from the web. Instead, it constructs tasks by reverse-engineering the hyperlink structure of a web corpus.

  • The Analogy: Think of it like a game of "Six Degrees of Kevin Bacon," but for facts. The authors start with a seed page and follow its out-links to build an "entity graph" of connected facts.
  • Making it Hard: They rewrite every non-answer entity into a descriptive reference. For example, instead of saying "Paris," they might say "the capital of France, located in the region of Île-de-France." This prevents the agent from simply matching strings (pattern matching) and forces it to actually reason through the connections.
  • The Dual Filter: They then use a "reference model" to filter questions. They keep only questions that the model cannot answer from memory (closed-book) but can answer if the relevant facts are provided (open-book). This ensures the questions are challenging yet solvable.

2. Training the Agent (SFT & RL) Once the hard questions are built, the authors train two models: Iris-mini (35B parameters) and Iris-pro (397B parameters).

  • Supervised Fine-Tuning (SFT): A "teacher" model solves these constructed questions using a ReAct loop (Reason + Act). The resulting successful trajectories (sequences of reasoning and tool use) are collected. However, not all trajectories are good—they might be degenerate, repetitive, or too short (just a direct lookup). The authors filter these aggressively at both the trajectory level and the individual turn level using an LLM judge. What's left is high-quality training data.
  • Reinforcement Learning (RL): The policy is then optimized against live search. This is crucial: the agent actually interacts with the web. To make this efficient, they use a "partial rollout" strategy. If a search session is taking too long, they don't start from scratch; they save the progress (the "prefix") and resume it later. They also run several internal "Qwen3.5-397B" models inside their training cluster to act as the judge and summarizer, eliminating the need for slow external API calls.

3. The "Climbing" Loop The magic happens when they alternate these two stages. After an RL round, the best-performing rollouts (the most efficient successful searches) are distilled back into the policy via SFT. The policy improves, and the next RL round targets harder questions. This creates a "curriculum" that automatically gets harder as the model gets smarter.

Key Results: Numbers That Translate to Impact

The paper reports results on four benchmarks, comparing their two models against other open-source agents. Because context management (CM) significantly affects scores, the results are presented both ways.

With Context Management (the "discard-all" strategy enabled):

  • Iris-mini (35B): Achieves 82.2 on BrowseComp, 84.8 on BrowseComp-ZH, 86.9 on DeepSearchQA, and 52.3 on HLE.
  • Iris-pro (397B): Achieves 88.6 on BrowseComp, 85.1 on BrowseComp-ZH, 92.9 on DeepSearchQA, and 56.4 on HLE.

Translating these numbers:

  • On DeepSearchQA, Iris-pro scores 92.9. The paper notes this means the model successfully recovers nearly all the required evidence for deep research questions. Compared to the previous best open-source model (XYZ-Aquila-mini at 86.9), Iris-pro improves the F1 score by roughly 6.4 percentage points. In plain language: if the test had 100 complex questions requiring multiple pieces of evidence, Iris-pro would get about 6 more questions fully right than the previous best model.
  • On BrowseComp, Iris-pro scores 88.6, beating the strongest 30B model by 3.8 points. This signifies a significant leap in the model's ability to navigate long-tail entities and indirect clues.

The "No Context Management" Baseline A crucial part of the paper is showing that these models are genuinely capable, not just relying on a trick to reset context. Without any CM:

  • Iris-mini still scores 64.7 on BrowseComp and 72.3 on BrowseComp-ZH, outperforming many previous state-of-the-art models that did use CM.
  • Iris-pro scores 72.6 on BrowseComp and 76.8 on BrowseComp-ZH.

This proves that the "search intelligence" is built into the model's weights, not just the software surrounding it.

Why It Matters: Key Takeaways

  1. Search as an "Atomic Capability": The authors argue that search ability is a fundamental skill, not just a "vertical" for web browsing. Their models, trained specifically for search, transfer positively to other domains like general tool use and office workflows (OfficeQA, APEX). This suggests that training an AI to browse the web effectively makes it better at acting under incomplete information in general.
  2. The CM Trade-off: Context Management is a powerful tool. The paper shows that for the smaller model (Iris-mini), CM is worth up to 21.2 points on BrowseComp. This is because the smaller model takes more steps to find answers, hitting the context limit more often. For the larger model (Iris-pro), the gains are smaller (around 16 points) because it is more efficient with its steps. This highlights that CM is most valuable when the model has learned good behaviors but lacks the "working memory" to hold the whole search history.
  3. Open-Source Progress: The models achieve the "strongest overall results among open-source search agents in their respective parameter ranges." This is significant because it means a 35B model can compete with much larger, heavier-compute systems from other labs, and a 400B model can push the frontier of what open-source agents can do.
  4. The "Inconsistency" Warning: The paper ends with a note on benchmark quality. In one case (BrowseComp-ZH Q85), the agent gave the logically correct answer ("Bolton") based on the series' events, but the official ground truth was "Lannister." This highlights that even benchmark annotations can have inconsistencies with source material, and the community needs to be vigilant about answer keys.

In summary: Iris-mini and Iris-pro represent a significant step toward reliable, general-purpose search agents. By rigorously constructing hard questions, filtering noisy trajectories, and iterating between supervised and reinforcement learning—all while transparently reporting performance with and without context management—the authors provide a strong open-source baseline and a recipe that the broader research community can potentially replicate or build upon.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →