arXiv:2609.04482nvidia/nemotron-3.5-lightning-30b-a3bSeptember 3, 2026

Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety RefusalExplained for Beginners

Alejo López-Ávila, Iker García-Ferrero, Jezabel Garcia +2 more

Computation and Language

Abstract

Safety alignment is usually posed as a topic-level question: is this subject harmful? Deployments ask a narrower one. A civics tutor and a public-sector assistant may share a base model yet need different boundaries inside the same topic, refusing targeted political manipulation while still answering factual questions about the same election. We formulate this as narrow-boundary safety and introduce an offline self-generated framework combining controlled topic generation, coverage repair, in-distribution compensation data, and harmful-benign pairs for training and evaluation. Single-shot generation leaves 19.88% of prompts without accepted refusal traces, whereas escalating retries leave 0.20%. On political persuasion with Qwen3-8B, training on refusal data completed through Escalate increases target-domain refusal from 9.47% to 84.75% and reduces the mean unsafe-response rate across three broader harmfulness benchmarks from 26.26% to 0.14%, but increases XSTest over-refusal from 2.00% to 74.00%. In a separate matched comparison, replacing external responses with verified target-model responses reduces over-refusal from 15.20% to 5.20%. Boundary-pair data reduces comply-side over-refusal on held-out pairs from 32.94% to 4.16%, while harmful-side refusal decreases only from 91.88% to 87.72%. These results show that data composition controls the safety and usability trade-off, and that safety alignment should be evaluated on both sides of the intended refusal boundary.

The Problem: Refusal Beyond "Topic"

Imagine two different apps built on the same large language model. One is a civics tutor helping students understand an upcoming election. The other is a public-sector assistant helping citizens fill out forms related to that same election. Both need the model to be "safe," but they need very different boundaries.

The civics tutor might want the model to refuse requests asking for persuasive arguments about a specific candidate—essentially refusing political manipulation. The public-sector assistant, however, needs to answer factual questions about election procedures, voting laws, and candidate histories.

Current safety alignment usually treats this as a "topic-level" question: Is this subject harmful? If the topic is "politics," the model is often tuned to refuse almost everything related to it. But real-world deployments rarely want a blanket refusal of an entire subject. They want to refuse only the harmful subset—like targeted manipulation or propaganda—while still answering the factual, benign questions.

This paper calls this "narrow-boundary safety." The core problem is that existing methods are too blunt. They either refuse everything on a topic (hurting usability) or refuse nothing (failing safety). The authors argue that safety should be evaluated on both sides of the intended refusal boundary: the model should refuse the bad stuff, but it should not refuse the good stuff that happens to share surface similarities.

How It Works: The Self-Generated "Safety Recipe"

The authors aren't collecting human-written safety examples from scratch. Instead, they’ve built an offline pipeline that generates its own safety data from the model being trained. Think of it as the model giving itself a homework assignment and then grading its own answers.

Here is the step-by-step technical mechanics, explained with an analogy:

  1. Controlled Topic Generation (The Syllabus): The pipeline starts with a broad topic (e.g., "Politics"). It then hierarchically breaks this down into subtopics, specific intents, and even specific "personas" (character roles the user might adopt). It also controls the length and style of the prompts. This ensures the data covers the full spectrum of edge cases, from short, snappy queries to long, detailed requests.
  2. Single-Shot Generation (The First Attempt): The model is prompted to respond. For harmful prompts, it’s given a "refusal steering" instruction. For benign prompts, it’s left alone. The system checks if the model actually refused.
    • The Analogy: Imagine a teacher assigning a worksheet and collecting it after one try. Many students might not have finished or understood the task.
  3. The Coverage Gap (The Dropout Problem): The authors found that in this "single-shot" attempt, 19.88% of prompts didn't get an accepted refusal trace. These prompts were simply dropped from the training data. The authors argue this is dangerous because the prompts that fail are often the hardest, most edge-case examples. If you drop them, the model never learns to handle them.
  4. Coverage Repair (The Retry): To fix this, the pipeline uses two strategies:
    • Graft: If a prompt failed, the system pairs it with a pre-written, neutral refusal snippet. It’s like giving a student a model answer when they got stuck.
    • Escalate: The system retries the prompt multiple times, using progressively stronger "refusal instructions" each time (like a teacher rephrasing a question until the student understands).
    • Escalate + Graft: A combination of both.
    • With these retries, the drop rate plummets to just 0.20%. The dataset grows from roughly 32,000 prompts to over 40,000, ensuring almost every edge case is captured.
  5. Boundary Pairs (The Local Test): This is the cleverest part. For every harmful prompt the model learns to refuse, the system creates a "twin" benign prompt. They share the same topic anchor (e.g., "elections") but differ in intent (one asks for persuasion/refusal, the other asks for factual explanation/compliance).
    • The Analogy: It’s like having a "Good Cop / Bad Cop" pair of questions. The model is tested on whether it refuses the "Bad Cop" (manipulation) but answers the "Good Cop" (facts).
  6. Compensation Data (The Guardrails): The system also creates specific data to prevent two failure modes:
    • Fake-Harm: Benign prompts that look harmful on the surface (e.g., the word "attack" in a political context) but are actually safe. The model is trained to not refuse these.
    • In-Distribution Compliance: Instead of using pre-written "compliant" answers from other models (which can make the model sound robotic or shift its personality), the system uses verified responses from the model itself. This keeps the model's natural "voice" while teaching it when to say "no."

Key Results & Benchmarks: The Trade-Off

The results are striking, but they highlight the central tension the authors identified: You cannot improve safety without affecting usability.

  • The Safety Win: Training on this carefully constructed refusal data (specifically using the "Escalate" strategy) dramatically increased the model's ability to refuse target-harmful prompts. On political persuasion with the Qwen3-8B model, the target-domain refusal rate jumped from a meager 9.47% to a strong 84.75%. Furthermore, the model's overall unsafe behavior on broader benchmarks dropped from 26.26% to 0.14%.
  • The Usability Cost (Over-Refusal): This is the catch. Because the model is being trained to be very strict about refusal, it started refusing things it shouldn't. On the XSTest benchmark (which tests if models over-refuse safe prompts), the over-refusal rate skyrocketed from 2.00% to 74.00%.
    • The authors describe this as the model becoming a "blunt refusal machine." It’s safe, but it’s useless—it won't answer legitimate questions because it’s too scared of getting something wrong.
  • The Fix: Boundary Pairs & Verified Responses: The paper shows that the right data composition can mitigate this. When the authors added "Boundary-pair data" (the twin prompts mentioned earlier), they successfully reduced the over-refusal on the "comply side" from 32.94% down to 4.16%. This means the model learned to distinguish between the harmful manipulation prompts and the safe factual prompts much more accurately.
  • The Verification Boost: Interestingly, when the researchers replaced external compliance responses with verified responses from the model itself, the over-refusal dropped even further, from 15.20% down to 5.20%. This proves that keeping the model's own "voice" matters for maintaining usability.

Why It Matters: Key Takeaways

This paper is a significant step toward making AI assistants actually useful in specific professional domains without sacrificing safety. Here are the four key takeaways:

  1. Safety is a Boundary, Not a Wall: You shouldn't have to choose between a model that refuses everything related to a topic (like all of politics) and a model that refuses nothing. The goal is to learn a precise boundary within the topic. This paper provides the framework and data to train for that boundary.
  2. Data Composition is the Knob: The authors show that you can heavily influence the safety-usability trade-off simply by what you include in the training data. Adding "boundary pairs" (harmful vs. benign twins) and "verified in-distribution responses" are specific levers that shift the balance toward usability without sacrificing safety.
  3. The Danger of the "Drop-Set": The finding that 19.88% of prompts are silently dropped in single-shot generation is a critical insight for anyone building safety systems. It suggests that standard self-generation methods are leaving huge gaps in the model's safety training, particularly on difficult edge cases.
  4. Evaluate Both Sides: The paper’s most important methodological contribution is the insistence on evaluating safety on both sides of the boundary. A model might look "safe" because it refuses harmful prompts, but if it also refuses 74% of safe prompts (as seen in the XSTest results), it has failed as a useful tool. Future evaluation must measure this trade-off simultaneously.

Summary

Safety for Whom? argues that we need to move away from "topic-level" refusal toward "boundary-aware" refusal. By using a self-generated, offline pipeline that repairs coverage gaps and carefully constructs harmful/benign pairs, the authors successfully trained a model to refuse targeted political manipulation while still answering factual questions. The results demonstrate that with the right data composition, it is possible to achieve high refusal rates on harmful content while keeping over-refusal low, making the model both safe and usable for specific deployments.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →