What Else Needs Fixing? Exploring Cost-Effective Test-Time Compute for Revision Propagation in Artifacts Generated Through ConversationExplained for Beginners
Daisuke Kikuta
Abstract
Large Language Models (LLMs) often help users generate artifacts through iterative cycles of generation and revision in conversation. A challenge here is that, when users specify only a local change during revision, LLMs must instead identify the relevant dependencies and propagate the revision to all affected parts of the artifact. This paper studies this ability of LLMs on conversationally generated artifacts, where the artifact context and its dependencies may be embedded in the conversation history. Toward practical use, we also explore cost-effective test-time compute for this new setting. Specifically, we introduce a new benchmark for this setting, and evaluate nine revision methods, including sequential reflection and parallel sampling variants, using gpt-oss-20b/120b, gpt-5.4-mini, and qwen3.5-9b/27b/122b on the benchmark. The results show that baselines achieve accuracies of 68.3--93%, and the most cost-effective method is selecting from three parallel samples using either LLM-based or medoid selection, which improves accuracy by 2.2--9.7%. Our code and dataset are available at https://github.com/ntt-dkiku/llm-revision-propagation.
1. The Problem
Imagine you are chatting with an AI to plan a complex trip. You start with a general idea—"I want to visit the Statue of Liberty, then maybe a museum, and catch a show in the evening." The AI generates a JSON itinerary for you. A few turns later, you realize, "Actually, start the trip two days later." You give that single, local instruction. The AI now faces a difficult task: it must look at your original request, understand how all the dates in the itinerary depend on the start date, and automatically shift every relevant day forward by two days. If it misses just one dependent date—say, the dinner reservation—the artifact becomes inconsistent, and the user has to do manual cleanup.
This paper identifies that exact gap. Large Language Models are increasingly used to generate artifacts—documents, code, travel plans—through iterative conversation. The prevailing assumption has been that users will spell out every change, or that the AI can rely on explicit dependency graphs (like import lists in code or citation networks in papers). But in natural conversation, dependencies are often implicit. They live in the chatter: "three days after the trip begins," or "the meeting the week after the conference."
When a user specifies only a local change, the LLM must act as a detective. It has to identify which parts of the artifact are silently tied to the changed element and propagate the revision across the whole structure. This problem—revision propagation in conversationally generated artifacts—has been studied before, but usually in settings where dependencies are explicit (like code repositories with their call graphs). This paper argues that real-world, conversational artifact generation is a different beast: the dependencies are often hidden in the conversation history itself, not in the final JSON.
The real-world significance is straightforward. As we offload more planning, coding, and document drafting to AI, the "local edit" becomes the primary mode of interaction. If the AI can’t automatically propagate those edits, the productivity gain vanishes, and the user ends up doing the AI's job. This paper sets out to measure how well current models handle this, and how to do it without breaking the bank.
2. How It Works (The Technical Mechanics)
The paper introduces RevPropBench, a benchmark designed to test this specific ability. The setup has two phases. First, an LLM and a user "build" a JSON artifact through a multi-turn conversation. The AI adds elements one by one—dates, items, configuration settings—until a final artifact is formed. Then, in the revision phase, the user gives a local instruction (e.g., "change the start date"), and the LLM must output a JSON Patch (following the RFC 6902 standard) that updates not just the targeted element, but every other element that depends on it.
The benchmark is cleverly constructed. The authors synthesize conversations and then have humans annotate the "gold patches"—the correct set of changes the model should make. This ensures the tasks are realistic and the dependencies are truly implicit.
The Technical Mechanics: A Software Engineer’s Analogy
To understand the core method, it helps to think like a software engineer. Imagine the JSON artifact is a runtime object graph. In code, if you change a function's input parameter, you have to trace every function that calls it, every variable that parameter feeds into, and every output that changes as a result. It’s a web of dependencies.
The paper evaluates several "revision methods"—strategies the LLM uses to figure out that web.
- Baselines (Single Pass): The simplest approach. The model just looks at the current JSON state (J), or the conversation history (H), or both (J+H) and generates a patch in one go. The paper finds that J+H is the strongest baseline because providing both the conversation context and the current state gives the model the best of both worlds: it remembers how things were built and what the current structure looks like.
- Sequential Reflection (REFLECT): This is akin to the "Self-Critique" pattern. The model generates an initial patch, "sees" the result by applying it to the artifact, and then iteratively fixes any errors it finds. It’s like a developer writing code, compiling it, seeing the error message, and fixing the bug. The paper notes this reduces "misses" (forgotten updates), but the gains are modest (+0.5–8.2%).
- Parallel Sampling with Medoid Selection (MED): This is the "wisdom of crowds" approach. The model generates, say, five different patches in parallel. Instead of trying to merge them all together (which can get messy), it picks the single best one. The "medoid" is the patch whose changes are the most "typical" or central compared to the others. If four out of five patches agree that a certain date should move, but one outlier says "leave it alone," the medoid picks the majority view. It’s a way to filter out noise.
- Parallel Sampling with LLM-based Selection (SELECT): This is perhaps the most interesting method for product managers. The model generates multiple patches in parallel, and then—here is the key—it uses the same model (but a separate "thinking" step) to act as a judge and pick the best patch. The paper frames this as a cost-effective way to get the benefits of sampling without needing a separate, smaller verifier model.
The Analogy: The Group Project
Think of SELECT like a group project presentation. Five team members (the parallel samples) each draft a section of the revised itinerary. Then, a "selector" (the LLM judge) reviews all five drafts and picks the one that looks most complete and accurate, striking a balance between following the user's local request and respecting the overall trip plan. It avoids the deadlock of "AND" (where everyone must agree, which is impossible) and the chaos of "OR" (where anyone’s idea is accepted, leading to errors).
3. Key Results & Benchmarks
The results are the heart of the paper. Across six different models (ranging from a compact 9-billion parameter Qwen to a massive 122-billion parameter Qwen, and including OpenAI’s GPT-5.4-mini), the baseline accuracy (using the J+H approach) ranges from 68.3% to 93%.
Here is the translation of those numbers into plain language:
- The "Miss" Problem: The paper’s failure analysis reveals that the most common error type is "miss"—the model simply forgets to propagate the revision to some dependent element. Models tend to propagate too little rather than too much.
- The SELECT Boost: The paper’s central finding is that SELECT (selecting from three parallel samples) is the most consistent and cost-effective improvement. Compared to the single-pass baseline (J+H), SELECT improves accuracy by 2.2% to 9.7%.
- Translation: If a model was getting 80% of revisions exactly right, SELECT could push that to roughly 82-83% without needing a smarter model—just by thinking a little harder during inference.
- Model Scale Matters: Generally, larger models perform better. The hierarchy roughly follows: Qwen 9B < 27B < 122B and GPT-OSS 20B < 120B < GPT-5.4-mini. Notably, GPT-5.4-mini already hits 93% on the baseline, which the authors acknowledge as a "ceiling" that makes future comparisons tricky.
- The Medoid Alternative: MED (medoid selection) is a close runner-up. It improves accuracy by 1.8% to 7.7% and is notably cheaper and faster than SELECT because it doesn't require a separate "judge" LLM call to make the final choice.
4. Why It Matters (Key Takeaways)
The authors distill their findings into four key takeaways for practitioners:
- Conversation History is Gold: The baseline results consistently show that providing the conversation history (H) alongside the final artifact (J) is crucial. The history contains contextual clues about dependencies that the final JSON alone cannot convey. If you are building a product that lets users revise AI-generated artifacts, never discard the chat history.
- Parallel Sampling is the Low-Hanging Fruit: For most models, the cheapest way to get a reliability boost is SELECT (picking the best of three parallel tries). It offers the best trade-off between accuracy gain and computational cost. If you need a quick win in a production system, enable this mode.
- Meddle with Medoids for Speed: If latency is your primary concern (i.e., you can't afford the extra time it takes for the model to generate multiple samples and then judge them), MED is the alternative. It provides most of the accuracy benefit of SELECT but runs faster because it doesn't need that second "selection" stage.
- The "Ceiling" Warning: The fact that GPT-5.4-mini scores 93% on the baseline suggests that as models get smarter, bare-minimum prompting will become less likely to fail. However, the paper’s benchmark and methods remain valuable because even smart models make "silly" mistakes—like missing a single dependent date in a travel plan—and the test-time compute methods (SELECT/MED) are the safety nets that catch those edge cases.
Summary This paper tackles the frustrating scenario where a user makes a small, local change to an AI-generated artifact, and the AI fails to spot the ripple effects. By creating a rigorous benchmark (RevPropBench) and testing various "test-time compute" strategies, the authors demonstrate that simply asking the model to generate a few parallel options and then picking the best one (SELECT) is a highly effective, cost-free way to improve reliability. For anyone building AI-assisted workflows involving iterative revision, the key insight is clear: keep the conversation history, and use parallel sampling with medoid or LLM-selection to ensure nothing falls through the cracks.
Want to understand AI papers like this from scratch?
Follow the free AI Learning Roadmap →