arXiv:2609.04061nvidia/nemotron-3.5-lightning-30b-a3bSeptember 3, 2026

When Models Edit Too Much: On the Fidelity of Minimal Code EditsExplained for Beginners

Tongyao Zhu, Wei Hern Lim, Min-Yen Kan

Software EngineeringArtificial IntelligenceComputation and Language

Abstract

Large language models (LLMs) are increasingly used to edit existing code, but correctness alone is not enough: useful repairs should also be minimal, reviewable, and faithful to the original implementation. We study over-editing, the tendency of a model to rewrite code beyond what is required to fix a bug. We construct an evaluation framework from 400 BigCodeBench problems by injecting controlled AST-level corruptions into reference solutions, giving each repair task a known minimal patch. Across frontier LLMs, over-editing is widespread even among strong models like GPT-5.5: high Pass@1 can coexist with unnecessarily large edits and added cognitive complexity. A preservation instruction substantially reduces this behavior, lowering average excess Levenshtein distance from 0.195 to 0.131, reducing added cognitive complexity by 26.6%, and increasing Pass@1 by 2.3 points. However, these gains do not simply follow from a larger reasoning budget or larger models. We next ask whether minimal editing can be learned directly during post-training. We observe that supervised fine-tuning overfits to seen corruption patterns, whereas reinforcement learning gives the best out-of-domain edit-fidelity and performance-retention trade-off. These results position edit fidelity as a distinct axis of code-repair quality and show that it can be measured and learned.

Here is a structured explanation of the paper "When Models Edit Too Much: On the Fidelity of Minimal Code Edits."

1. The Problem

The paper identifies a specific failure mode in how Large Language Models (LLMs) handle code repair. When developers use LLMs to fix bugs in existing code—brownfield maintenance—they expect a minimal, targeted patch. However, the paper argues that LLMs frequently exhibit over-editing: they rewrite more code than is strictly necessary to fix the bug.

The real-world impact of this is significant. In professional software development, a correct but large rewrite increases the "review burden." A reviewer must sift through unnecessary lines to understand what actually changed, which raises the risk that a subtle regression (a new bug) outside the existing test suite will go unnoticed. The paper emphasizes that correctness alone is not enough; the patch must be faithful to the original implementation.

To study this, the authors constructed a controlled evaluation framework. They took 400 problems from the BigCodeBench benchmark and injected known, controlled corruptions (bugs) into the reference solutions. Because the bug was injected by design, the "minimal fix" was known by construction. This allowed them to measure exactly how much extra code a model wrote compared to the ideal, minimal patch.

2. How It Works (The Technical Mechanics)

The core technical mechanism the paper evaluates is the model's ability to perform minimal token-level edits to reverse a injected bug.

The Metrics:

  • Excess Levenshtein Distance: This measures the "bloat" of the edit. If the minimal fix requires changing 2 tokens, and the model changes 20, the excess distance is high. A value of zero means the model made exactly the minimal change.
  • Added Cognitive Complexity: This measures structural overhead. Even if a patch is short, it might add if statements or nested loops that make the code harder for a human to review. The metric compares the cognitive complexity of the model's output versus the gold minimal repair.

An Analogy for Software Engineers: Think of a function as a carefully maintained Rube Goldberg machine designed to pop a balloon. Someone introduces a bug by taping the balloon to the ceiling instead of the floor. The minimal fix is simply cutting the tape (a one-line change). However, an over-editing model might decide to redesign the entire machine, moving furniture, rewiring circuits, and painting the walls—all while successfully popping the balloon. From a test perspective, the goal is achieved, but the "repair" is massive and fragile.

The paper finds that this behavior is widespread. Even top-tier models like GPT-5.5 achieve high Pass@1 (the percentage of outputs that pass the tests) while making edits that are 10x to 30x larger than the minimal required change. The default behavior of these models appears geared toward "producing a robust solution" rather than "recovering the smallest local patch."

3. Key Results & Benchmarks

The paper presents several striking quantitative findings across frontier models:

  • The Scale of Over-Editing: Among models whose outputs all pass tests, the median excess patch size ranges from 2 to 60 inserted lines. For example, GPT-5.4 adds 60 lines of input validation and dtype coercion for a bug that requires only a single line correction.
  • Prompting Helps: Simply appending a instruction to "keep as much of the original code as possible" substantially reduces over-editing. This lowers the average excess Levenshtein distance from 0.195 to 0.131 (a 33% reduction), reduces added cognitive complexity by 26.6%, and surprisingly increases Pass@1 by 2.3 points. This suggests that the "preservation" instruction changes the model's latent mode, treating the existing code as evidence to preserve rather than a reference to improve upon.
  • Reasoning is Not a Silver Bullet: The paper tests "reasoning" variants (like chain-of-thought) and finds they do not consistently reduce over-editing. In some cases, reasoning models actually produce more cognitive complexity. The pattern is highly model-specific.
  • Scale Does Not Fix It: Increasing model size (e.g., moving from 7B to 32B parameters) improves Pass@1 (the model stops failing to fix the bug), but it does not monotonically reduce excess edit distance. A larger model is still capable of making a 60-line edit when a 1-line fix is needed.

4. Why It Matters (Key Takeaways)

The authors conclude that edit fidelity is a distinct axis of code-repair quality. Here are the four key takeaways:

  1. Over-editing is steerable, not inevitable. Explicit preservation prompts are a low-cost, effective intervention. For the heaviest over-editors (like GPT-5.5), this approach nearly halves the excess distance. This means developers can improve repair quality without waiting for a new model version.
  2. Larger models and more reasoning do not guarantee minimal edits. A model can be highly capable at solving the problem functionally yet remain "bloated" in its implementation. Evaluation metrics must include edit-size measures alongside Pass@1.
  3. Reinforcement Learning (RL) is the most promising path for durable improvement. The paper shows that post-training a model using RL (specifically a GRPO-style objective that rewards passing tests and small edits) leads to the best out-of-domain edit fidelity. Crucially, RL preserves broader coding ability better than Supervised Fine-Tuning (SFT), which tends to overfit to the specific corruption patterns it was trained on.
  4. The "Granularity Mismatch". The paper observes that models often identify the bug correctly but treat the surrounding code as suspect. For instance, a one-token off-by-one error might trigger the model to add defensive checks for empty lists or change data structures. The model behaves as if asked to deliver robust, production-grade code, whereas the evaluation measures whether it recovers the original intent.

Summary The paper reveals that our current benchmarks for code LLMs are "blind" to a common failure mode: passing the tests while breaking the code's readability and maintainability. By constructing a ground-truth framework of minimal repairs, the authors show that we can measure and improve "edit fidelity." The most immediate fix is a simple prompt tweak, but for long-term robustness, post-training with reinforcement learning appears to be the most effective way to teach models to keep their edits minimal.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →