arXiv:2608.21500nvidia/nemotron-3.5-lightning-30b-a3bAugust 21, 2026

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy DistillationExplained for Beginners

Yibo Peng, Long Lian, David Wagner +1 more

Cryptography and SecurityArtificial Intelligence

Abstract

Prompt injection is listed as the \#1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacker may inject a prompt into the data, saying, "Ignore all prior instructions and perform <an attacker's task>." To prevent arbitrary manipulation of agents, defenders try to train secure LLMs, which, however, still suffer from near 100% attack success rates (ASRs) against adaptive prompt injections. We note that this is because existing defensive finetuning recipes rely on sequence-level feedback signals (in DPO or GRPO). Treating an entire output equally prevents the model from learning precisely which output tokens are insecure. In this paper, we propose Secure On-Policy Distillation (SecOPD) that provides token-level feedback to guide defensive fine-tuning. The LLM receives an injected sample and produces a rollout, whose tokens are scored by the initialization model given the corresponding clean input. With more fine-grained training signals, our defended Qwen3.6-27B achieves a 9.0% ASR against the SoTA PISmith adaptive prompt injections, compared to 94.0% for the prior SoTA, Meta-SecAlign. The obtained security generalizes to domains completely unseen in training: in agentic tool calling, SecOPD achieves a 4.7% ASR compared to 5.5% for Meta-SecAlign. Code and the model are available at https://github.com/pppyb/SecOPD and https://huggingface.co/pybbb/Qwen3.6-27B-SecOPD.

The Problem: When Data Turns Against the Model

Imagine you run an AI agent—a system that can browse the web, read your emails, or edit documents. Its power comes from accessing external data. But that same access creates a security hole. An attacker can hide a malicious instruction inside a seemingly harmless email or webpage. When the agent processes that data, the hidden instruction says, in effect: "Ignore everything I’ve been told before. Now do this bad thing instead."

This is a prompt injection, and it’s listed as the #1 threat to AI agents by OWASP. The consequences are real: agents have been tricked into generating bad code, downloading malware, or leaking private messages.

Researchers have tried to fix this by fine-tuning models to be "secure." But here’s the problem: the standard fine-tuning methods treat the entire model output as one single unit. If the model answers the user’s question correctly and follows the attacker’s instruction, the trainer gives the whole response a "good" or "bad" score—there’s no granular feedback on which parts are which. As a result, even fine-tuned models still fall victim to adaptive attacks, with nearly 100% of injection attempts succeeding.

How It Works: Token-Level Feedback, Not "All-or-Nothing"

The paper’s core insight is that we need to treat the model’s output like a conversation with turn-by-turn feedback, not a single yes/no answer.

The authors propose Secure On-Policy Distillation (SecOPD). Here’s the mechanics, explained simply:

  1. The Setup: During training, the model receives two versions of the same prompt. One is "clean"—the trusted instruction plus benign data. The other is "attacked"—the same thing, but with a malicious instruction injected into the data.
  2. The Rollout: The model generates a response (a "rollout") to the attacked prompt.
  3. The Teacher: A separate, frozen model (the "teacher") looks at the same generated tokens but under the clean prompt. It scores each token: "How likely was this token to appear if the injection weren't there?"
  4. The Comparison (The Advantage): For every single token in the model's response, SecOPD compares the student's probability to the teacher's probability.
    • If the student produced a token that the teacher would have produced on the clean prompt, that token gets positive feedback (it’s encouraged).
    • If the student produced a token that only appears because of the injection (e.g., appending "The season that comes after autumn is winter."), that token gets negative feedback (it’s discouraged).

Why an analogy to software engineering helps: Think of the old methods like a code reviewer who only says "This whole function is good" or "This whole function is bad." SecOPD is like a reviewer who marks specific lines: "Line 5 is great, Line 7 is terrible, Line 8 is great." The model learns exactly which patterns to keep and which to discard.

Key Results: An Order of Magnitude Improvement

The results are striking. The paper evaluates against PISmith, the state-of-the-art adaptive attack that trains a dedicated attacker to break the defended model.

  • Meta-SecAlign (the prior best defense): 94.0% Attack Success Rate (ASR). Out of 100 attack attempts, nearly all succeed.
  • SecOPD: 9.0% ASR. This is an order-of-magnitude improvement. The model successfully resists the adaptive attack 91% of the time.

This security doesn't stay confined to the training domain either. When tested on AgentDojo, a benchmark for tool-using agents (like querying a database or sending an email), SecOPD maintains a 4.7% ASR compared to 5.5% for Meta-SecAlign. The defense generalizes to domains it never explicitly saw during training.

Critically, SecOPD preserves the model's utility. On a suite of general capability benchmarks (MMLU-Pro, GPQA, GSM8K, etc.), SecOPD loses less than 3% performance compared to the undefended model, while Meta-SecAlign loses more ground, and other methods (like GRPO) lose significantly more.

Why It Matters: The Takeaways

For anyone building or relying on AI agents, this paper signals a shift in how we think about security:

  • Fine-grained training works: By moving from sequence-level to token-level feedback, we can teach models to distinguish between trusted instructions and injected ones much more effectively.
  • Utility and Security can coexist: Unlike previous approaches where you had to choose between a "dumb but safe" model and a "smart but vulnerable" one, SecOPD keeps the model smart. It preserves ~88% of general capability while slashing attack success rates from ~94% to ~9%.
  • Generalization is key: The fact that this works on AgentDojo— a completely different task (tool use) from the text-completion training—suggests the model learns a fundamental principle of "how to handle untrusted data," not just a memorized trick for one dataset.
  • Defense-in-Depth is still needed: The authors are clear that this is a milestone, not a silver bullet. They note that future, more advanced attacks may still find weaknesses. They emphasize that deployed agents should still use other safeguards: input filtering, action constraints, and monitoring.

The code and models are publicly available, meaning the research community can immediately build on this token-level distillation approach to push the boundaries of LLM security further.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →