One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM RecommendersExplained for Beginners
Minghao Luo, Liang Chen
Abstract
Search-augmented LLMs increasingly mediate everyday consumer recommendations by retrieving live web content. This creates a new risk: LLM recommenders may consume web content that Generative Engine Optimization (GEO) operators have polluted to mislead them. We ask: to what extent do they become unwitting promoters of fake products? We introduce FORGE (Fake Online Recommendations in Generative Environments), which locally rewrites real products in a frozen set of retrieved web pages into fake ones and measures how often the LLM recommends the fake product, across 225 real products in 15 categories and 5 consumer scenarios. Across 12 commercial and open-weights LLMs, all models are vulnerable: a single polluted page yields fooled rates of up to 27%, while the full top-3 replacement raises this to 73.8%. Vulnerability varies across categories, increasing when models lack stable prior knowledge of the products. Reasoning does not mitigate this vulnerability; instead, it often generates spurious social proof to justify false recommendations. None of the four defenses is adequate: the skepticism prompt can exacerbate vulnerability much like reasoning, the two consensus filters risk suppressing legitimate products, and credibility re-ranking helps every model but removes only a sixth of the fakes. We release the FORGE benchmark and the evaluation code at https://github.com/leoluolol/forge-benchmark.
One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders
1. The Problem
Imagine you ask an AI assistant for the best new smartphone or a reliable plumber in your area. The assistant performs a web search, pulls in a few top results, and synthesizes an answer. But what if those search results have been quietly tweaked? Not by hacking the system, but by simply creating a few plausible-sounding user reviews or forum posts that mention a fake brand.
This is the core risk that the paper “One Polluted Page Is Enough” investigates. As search-augmented LLMs become the primary interface for everyday recommendations—from dinner plans to tech purchases—they introduce a new vulnerability. These models don’t just sit in a static knowledge vacuum; they reach out and consume live web content. If Generative Engine Optimization (GEO) operators can get a fake brand into those top search results, the LLM can become an unwitting promoter of a product that doesn’t even exist.
The authors introduce FORGE, a benchmark designed to measure exactly this: how often an LLM recommends a fake product after it has been locally “implanted” into a frozen set of retrieved web pages. The findings are striking. Across 12 commercial and open-weights models, all of them are vulnerable. A single polluted page at the very top of the search results can fool a model 27% of the time. If the attacker places polluted pages in the top three slots, that success rate jumps to 73.8%. The attack works across categories, though it is most effective where consumers rely on community taste rather than established brand names (like dining or personal services) and least effective where stable brand knowledge exists (like smartphones or PCs).
Perhaps most unnerving is what happens when the model actually follows the misinformation. The paper shows that “reasoning” or “thinking” often doesn’t save the user. Instead of rejecting the fake product, the model frequently generates spurious social proof—inventing fake community endorsements or technical details—to justify why the fake product is recommended. The model essentially hallucinates a reputation for a product that has none.
2. How It Works (The Technical Mechanics)
To understand the mechanics, it helps to look at the typical pipeline a search-augmented LLM uses. When you query the system, it follows a path: User Query → Live Web Search → Top-K Evidence Bundle → LLM Consumption → Ranked Recommendation. The risk enters at the second step. The authors frame this as a “web pollution” problem distinct from traditional adversarial SEO. Traditional SEO tries to boost a real competitor’s ranking; this attack promotes a brand that is entirely fake—something the model has never seen before.
The FORGE benchmark operates by taking a frozen set of real retrieved web pages and, crucially, locally rewriting the dominant real brand mention into a fake brand-product compound. They preserve the URL, the surrounding context, the length, and the style. The only thing that changes is the brand name. Because the alteration is local and surgical, any shift in the model’s recommendation can be attributed directly to that brand swap.
The authors define three “attack styles” to test realism:
- Entity Replacement: Swapping the brand name (the default attack).
- Passage Injection: Adding a fake-brand paragraph to an otherwise untouched document.
- Full Synthesis: Replacing the entire document body with a synthetic fake review.
The default “entity replacement” is sufficient to cause havoc, but full synthesis is the strongest attack.
Regarding defenses, the paper tests four mechanisms, and unfortunately, none are silver bullets:
- Skepticism Prompt: Instructing the model to be cautious actually backfires for closed-source models, increasing vulnerability by an average of 24 percentage points.
- Consensus Filters (Prior and Agreement): These try to check if other documents or models agree, but they risk suppressing legitimate products. The “agreement filter” requires a brand to appear in 4 of 10 documents, which cuts fake recommendations but also discards 52%–79% of real ones.
- Credibility Re-ranking: This moves pages from user-generated content (UGC) to editorial sources before the model reads them. It helps every model and removes about a sixth of the fakes, but leaves a significant number still fooled.
3. Key Results & Benchmarks
The quantitative results paint a vivid picture of vulnerability.
- Universal Vulnerability: Across 12 models and 225 products, fooled rates span from 13.3% to 73.8% under the top-3 replacement attack. Even a single rank-1 polluted page induces failures in 27% of the most vulnerable model's cells.
- The “Reasoning” Trap: A paired experiment showed that when reasoning is enabled, models are more likely to be fooled. The gap reaches 18 percentage points on one model (Qwen3.5-9B). The authors attribute this to the model “talking itself into” the fake brand rather than dismissing it.
- Social Proof Hallucination: In the cells where models were fooled, the output often contained social-proof markers—phrases like “frequently recommended in communities” or “price-performance king”—that were entirely absent from the actual polluted documents. The models invented credibility.
- Category Matters: The attack is most successful in “low-prior-knowledge” categories. For example, in the Dining category, models were fooled at rates up to 93.3%. In contrast, technical categories like Phone/PC saw rates near 0%. The paper quantifies this via a “cross-model agreement” metric: if models generally disagree on what the real brands for a product are, they are much easier to fool.
4. Why It Matters (Key Takeaways)
The broader significance of this work touches on safety, user trust, and the future of search:
- The Ease of Attack: A small number of polluted pages (as few as three) can saturate the vulnerability, meaning a motivated GEO operator doesn't need to take over the entire web, just a few well-placed posts. This lowers the barrier to entry for misinformation campaigns targeting AI assistants.
- Reasoning is Not a Cure: The finding that “thinking harder” often makes models more susceptible is counter-intuitive and important. It suggests that for this specific type of pollution, increased computation time correlates with deeper engagement with the misinformation, leading to more confident (but false) recommendations.
- The Social Proof Illusion: The models don't just parrot the fake brand; they dress it up. This makes the misinformation stickier and harder for a human user to spot, because the answer sounds polished and authoritative, even though it's built on fabricated evidence.
- Defense Gaps: The current toolkit for defending against this is inadequate. Credibility re-ranking (prioritizing editorial content) is the most promising of the bunch, but it only removes about 17% of the fakes. The consensus filters are too destructive, suppressing legitimate user voices, and the skepticism prompt actively harms closed-source models.
In summary, this paper reveals a crack in the foundation of the new AI-assisted search experience. It shows that the "ground truth" these models rely on is surprisingly fragile. For product managers and engineers, the takeaway is clear: relying on live web retrieval for recommendations is risky, and simply telling the model to "be skeptical" or "reason through it" won't fix the problem. The path forward likely requires stronger evidence-level interventions—essentially, making the retrieval system more robust to injection before the LLM even sees the text.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →