arXiv:2608.22856nvidia/nemotron-3.5-lightning-30b-a3bAugust 24, 2026

Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QAExplained for Beginners

Jingjie Ning, Xueqi Li

Information RetrievalComputation and Language

Abstract

A retrieval-augmented QA system can return different answers after an index expansion even when its requested model identifier, prompt, retrieval policy, evidence depth, rendering, and exposed generation controls are held fixed. Aggregate accuracy may hide these changes when gains and losses cancel, while ordinary generation variability makes one-shot comparisons overstate update effects. We call the hidden phenomenon accuracy-blind answer churn and introduce the Snapshot Compatibility Audit, which estimates excess answer churn by subtracting same-snapshot repeat disagreement from cross-snapshot disagreement. We instantiate it by expanding one frozen FineWeb prefix from one to seven shards. In a preregistered 400-question Natural Questions study, normalized-exact and blinded-semantic excess churn are 6.44 and 10.25 percentage points while exact-match accuracy changes by only -1.50 points. A post-hoc analysis finds repeat-stable semantic flips on 40/400 questions. A separately preregistered 200-question TriviaQA study yields smaller, directionally consistent excess churn while exact-match accuracy moves in the opposite direction. An outcome-blind post-hoc 100-question subset replication with a second DeepSeek generator and serving configuration finds 8.75 pp of semantic excess churn even as exact match rises by 3.00 percentage points. Answer-level compatibility can therefore fail without a conspicuous or consistently directed utility shift. Retrieval-augmented releases should audit compatibility alongside utility.

Here is the structured explanation of the paper "Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA," written for a smart, non-academic audience.

1. The Problem: The Invisible Answer Flip

Imagine you are building a search engine or a QA assistant. You’ve tweaked the system to allow it to search a larger "brain" (the corpus). You keep the AI model, the prompt, and all the settings exactly the same. But when you deploy the bigger brain, users start getting different answers to the same questions.

In a typical evaluation, you would look at "accuracy"—the percentage of questions the system gets right. If that number stays roughly the same, you might shrug and say, "Everything is working fine."

But this paper argues that accuracy is a blunt instrument. It can miss critical changes because gains and losses can cancel each other out. A system might get 1% more questions right overall while completely changing the answer to a specific, important subset of questions. For a product manager or engineer, this is dangerous: if answers feed into caches, automated workflows, or user trust, a "green" accuracy dashboard can mask a system that is behaving unpredictably.

The paper identifies this hidden phenomenon as accuracy-blind answer churn. The real question isn't just "is the system right?" but "did the system change its mind, and if so, on which questions?"

2. How It Works: The Snapshot Compatibility Audit

To solve the problem of hidden changes, the authors introduce the Snapshot Compatibility Audit. This is a method designed to measure "compatibility"—whether the new version of the system behaves like the old one—rather than just "utility"—whether it is generally right.

The Core Analogy: The Faithful Intern

Think of the AI system as a smart intern. You ask the intern a question twice in a row. Sometimes, the intern is a bit stochastic—they might rephrase the answer or double-check a fact. If you ask them twice and get two different strings, that’s just the intern being the intern; it doesn’t mean the intern “changed their mind” about the fact.

Now, imagine you give the intern a larger library of books to reference. The Snapshot Compatibility Audit works like this:

  1. Measure Normal Variation: First, you ask the intern the same question twice using the small library. You note how often they disagree just due to their natural randomness. Let's say they disagree 10% of the time just by chance.
  2. Measure Cross-Expansion Variation: Now, you expand the library to seven times the size and ask the same questions twice. If the intern now disagrees 17% of the time, that’s a 7 percentage point increase.
  3. The Audit: The audit subtracts the "normal noise" (10%) from the "expansion noise" (17%). The remaining 7 percentage points is the excess churn—the answer changes that are specifically caused by the bigger library, not just the intern's natural variability.

Technical Mechanics (Simplified) The paper uses two specific ways to measure this "disagreement":

  • Normalized-Exact: This is the strictest measure. If the model outputs "The Devastator" in one snapshot and "The Executor" in another, and these don't match after the benchmark's text cleaning, they are considered different.
  • Blinded Semantic: This is a "softer" measure. An AI judge (blind to which snapshot the answer came from) decides if the two answers are factually equivalent, even if the wording is different. This catches cases where the model changes the phrase but not the fact.

3. Key Results & Benchmarks: The Numbers

The authors ran three main studies, and the results are striking. They show that exact-match accuracy (the traditional metric) tells only part of the story.

Study 1: Natural Questions (NQ) - The Main Event

  • The Setup: 400 questions, expanding the corpus from 1 shard to 7 shards.
  • The Accuracy Shift: Exact-match accuracy changed by only -1.50 percentage points. (This is the "dashboard" number).
  • The Churn Shift:
    • Normalized-Exact Excess Churn: +6.44 pp
    • Blinded-Semantic Excess Churn: +10.25 pp

What this means: Even though the overall accuracy score barely budged (dropping a tiny bit), the system’s answers shifted significantly. The "exact match" metric missed most of the action because the model often changed its wording or selected a different but valid answer span. The semantic audit, which understands meaning, caught a massive 10.25 percentage point shift. In plain language: on 400 questions, the system effectively "changed its mind" or selected a different answer strategy for over 10% of questions purely due to the larger corpus, even though it didn't help or hurt the overall score much.

Study 2: TriviaQA - The Supportive Check

  • The Setup: 200 questions.
  • The Results: Smaller excess churn (3.00 pp exact, 2.125 pp semantic), but the direction is consistent. Interestingly, the exact-match accuracy moved in the opposite direction (+1.25 pp) while the churn was positive. This confirms the paper's thesis: utility and compatibility can move in opposite directions.

Study 3: The Post-Hoc Replication (DeepSeek V4-Pro)

  • The Setup: A subset of 100 questions, using a different generator configuration.
  • The Results: 8.75 pp of semantic excess churn, even as exact-match accuracy rose by 3.00 percentage points.
  • The Takeaway: This is the strongest evidence for the paper's core claim. You can have a scenario where the system gets more questions "right" by the strict string-match metric, but its actual behavior has drifted significantly. Answers are flipping or changing form in ways the strict metric doesn't see.

4. Why It Matters: Key Takeaways

The authors distill their findings into four key points for practitioners:

  • Compatibility ≠ Utility: A system update can preserve or even improve aggregate accuracy while making the system behaviorally incompatible with its previous self. This is risky for production systems where consistent behavior is often more valuable than a slight accuracy bump.
  • The "Cancel-Out" Effect: Gains and losses in answer correctness often cancel each other out in aggregate scores. A system could be getting some questions right that it previously got wrong, while simultaneously getting other questions wrong that it previously got right. The net result looks stable, but the user experience is fragmented.
  • Generation Variability is a Baseline: The audit accounts for the fact that AI models are stochastic. By comparing "repeat calls" (asking the same system twice) against "cross-snapshot" calls (asking different corpus sizes), the authors separate normal AI "jitter" from actual corpus-induced changes.
  • The Call for Audits: The paper concludes that retrieval-augmented releases should audit compatibility alongside utility. Just as you wouldn't ship a software update without regression testing, RAG systems shouldn't be deployed with only an accuracy score. They recommend checking "answer-level compatibility" to ensure that expanding the knowledge base doesn't silently break the system's established behavior.

Summary

This paper reveals a "hidden curriculum" in Retrieval-Augmented Generation: the corpus you give the model matters, not just for accuracy, but for consistency. By using a rigorous audit method, the authors demonstrate that expanding a knowledge base can cause significant answer churn—shifting which answers the model picks—without moving the needle on traditional accuracy metrics. For anyone deploying RAG systems, this serves as a warning: a high accuracy score does not guarantee a stable or predictable product.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →