Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR ErrorsExplained for Beginners
Zhenghua Bao
Abstract
Speech-based applications pass spoken queries through automatic speech recognition (ASR) before any retrieval module, so ASR errors enter the pipeline as a fixed upstream constraint. We empirically test whether two extensions to standard retrieval-augmented generation (RAG), entity-graph linking and iterative reformulation, absorb or amplify these errors. Using four English accents synthesized through neural TTS, we evaluate four RAG configurations on three multi-hop QA benchmarks (HotpotQA, 2WikiMultiHopQA and MuSiQue) against a clean-text oracle. Although the structurally richer configurations generally retain higher absolute F1 under ASR input, both extensions amplify the error: the F1 gap from clean text to the highest-WER accent is 36-67% larger under their combination than under naive dense retrieval, on all three benchmarks. The dominant failure mode is corruption of one or more query entities, accounting for 87-96% of degradation cases on 2WikiMultiHopQA across all four methods. Two lightweight surface-form mitigations leave most of the gap intact, indicating that downstream retrieval structure amplifies remaining entity errors. We release code and data at https://github.com/ZhenghuaBao/spoken-multihop-rag .
Here is the explanation, structured according to your specifications.
1. The Problem
Voice-activated assistants and multimodal agents have become ubiquitous, but there is a hidden bottleneck in their pipelines: Automatic Speech Recognition (ASR). In these systems, a user’s spoken query is first transcribed into text by an ASR model, and only then does the retrieval-augmented generation (RAG) system kick in to find relevant information and formulate an answer.
The problem, as this paper demonstrates, is that ASR errors act as a "fixed upstream constraint." Because the retrieval module receives whatever text the ASR outputs—warts and all—any mistake the recognizer makes is baked into the subsequent reasoning steps. The authors zero in on a particularly challenging scenario: multi-hop Question Answering (QA), where the model must retrieve and reason over evidence across multiple documents to answer a question. This is harder than single-hop QA because it often requires linking multiple named entities (people, places, dates) to build a chain of reasoning.
The authors ask a critical question: When we add structural complexity to RAG—such as entity-graph linking or iterative reformulation—do these methods make the system more robust to ASR errors, or do they inadvertently make things worse? Using four synthesized English accents (US, Indian, Filipino, and Nigerian) created via neural Text-to-Speech (TTS), they systematically test how ASR errors propagate through different RAG architectures. The "oracle" condition in their experiments bypasses speech entirely, feeding clean text to the system, which serves as a baseline to measure exactly how much performance drops when ASR errors are introduced.
2. How It Works (The Technical Mechanics)
The paper evaluates four distinct RAG configurations, ranging from a simple baseline to complex hybrids:
- Naive RAG: The baseline. It performs top-k dense retrieval over a passage index and concatenates the retrieved documents as context for the generator. It has no special handling for entities or iterative reasoning.
- HippoRAG2: Adds an entity-graph layer. It extracts entity-relation triples from the corpus, builds a knowledge graph linking documents through shared entities, and retrieves information using personalized PageRank seeded from the query entities.
- IRCoT+Naive: Wraps Naive RAG in a Chain-of-Thought (CoT) loop. At each step, the LLM generates a reasoning sentence and a follow-up search query, retrieves new context, and repeats this process up to three times.
- IRCoT+HippoRAG2: The "heavy hitter." It combines the iterative reformulation of IRCoT with the entity-graph linking of HippoRAG2. This configuration achieves the highest F1 scores on clean text but, as the title suggests, has the worst robustness to ASR errors.
The Analogy: The "Broken Telephone" with a Map
Imagine you are playing "Broken Telephone," but you are trying to solve a mystery. The first person (the ASR) whispers the query to the second person (the Retrieval Module). If the whisper is garbled, the listener might mishear "Hanuman Patal Vijay" as "Honourable Peter Vijay."
- Naive RAG is like a player who just listens to the garbled whisper and tries to find the story in a big library based on that messed-up name. If the name is slightly wrong, they might grab the wrong book, but they can still scan the pages.
- HippoRAG2 is a player who has studied a map (the knowledge graph) of the library. They know that "Vijay" is a key that unlocks a specific corridor. Even if the name is slightly wrong (e.g., "Patil" instead of "Patal"), the map might still guide them to the right general area because the entity "Vijay" is recognized and used to navigate the graph.
- IRCoT is a player who, upon getting a wrong answer, stops and asks for clarification, re-reads the question, and tries a new search based on what they just learned.
The Twist: The paper shows that having a "map" (HippoRAG2) or a "clarification loop" (IRCoT) doesn't guarantee safety. In fact, the paper argues that these structures can sometimes amplify the error. For instance, if the iterative reformulation step latches onto the wrong entity (e.g., "Honourable Peter Vijay"), it might abandon the correct signals the original retriever had, locking the system into a death spiral of wrong reasoning. The structural richness creates more "pressure points" for the error to lodge itself, making the final answer more wrong than if the system had just struggled with the raw text.
3. Key Results & Benchmarks
The results are striking and center on the "oracle gap"—the drop in performance from the clean-text oracle to the worst-accented condition (Nigerian, NG).
The Amplification Effect While all methods suffer when ASR errors are introduced, the structurally complex methods amplify the damage compared to the simple baseline.
- Naive RAG on 2WikiMultiHopQA: The F1 gap from clean text to the Nigerian accent is 0.043.
- IRCoT+HippoRAG2 (the most complex method) on the same benchmark: The gap balloons to 0.195.
This means the complex method’s gap is 353% larger (roughly a 4.5x amplification) than the naive method’s gap. On the MuSiQue benchmark, the amplification is even more severe: the gap grows from 0.043 (Naive) to 0.072 (Complex), a 67% increase.
The Dominant Failure Mode: Entity Corruption The paper categorizes exactly why the answers fail. They find that the vast majority of failures are due to query-entity corruption—the ASR model mangles a named entity in the user's query.
- On the 2WikiMultiHopQA benchmark, entity corruption accounts for 87–96% of all degradation cases across all four RAG methods.
- On HotpotQA, it accounts for 67–82%.
- On MuSiQue, it accounts for 54–78%.
The authors illustrate this with a case study: A question about the entity "Hanuman Patal Vijay." Under a US accent, the ASR might mildly corrupt it to "Hanuman Patil Vijay" (a single letter change). Under a Nigerian accent, it becomes "Honourable Peter Vijay" (a severe distortion).
Why the mitigations fail The authors test two lightweight "rescue" strategies:
- N-best Decoding: Running the ASR multiple times at different temperatures to find a better transcription.
- Phonetic Correction: Using a phonetic algorithm (Double Metaphone) to find the closest matching entity in the database.
Unfortunately, these barely move the needle. Phonetic correction, the best of the two, only closes about 11% of the performance gap on the hardest configuration. This indicates that once the entity is corrupted, the downstream retrieval structure (the graphs and iterative loops) is so sensitive that fixing the "surface form" of the name isn't enough to save the answer.
4. Why It Matters (Key Takeaways)
- Structural Complexity is a Double-Edged Sword: Adding entity-graph linking and iterative reformulation improves performance on clean audio, but it makes the system fragile when the audio is noisy or accented. Engineers must balance the desire for higher accuracy on good input against the risk of brittleness on bad input.
- Entity Corruption is the Achilles' Heel: No matter how sophisticated the retrieval architecture is, if the ASR mangles a key name or date, the system is likely to fail. The paper suggests that future work should focus on "entity-aware" ASR or robust entity linking before the retrieval step, rather than relying on the RAG pipeline to fix recognition errors.
- The Gap is Real-World Significant: A 67% larger performance gap isn't just a statistic; it means that for a model that might answer 100 questions correctly on clean text, the complex system might answer significantly fewer correctly when faced with a strong accent, narrowing the usability of the technology for diverse user groups.
- Mitigation Requires More Than Surface Fixes: Simple tricks like N-best decoding or phonetic correction are diagnostic tools that reveal the problem but don't solve it. Truly robust multi-hop RAG will need architectural changes that can tolerate or correct entity errors at a fundamental level, rather than relying on the surface form of the query.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →