# Hugging Face — Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

- Company: Hugging Face (huggingface.co)
- Announced: 2026-09-08T14:23:07+00:00
- Category: research-paper
- Coverage: not counted
- Announcement: yes
- Group: announcements
- Source: https://huggingface.co/blog/MultiverseComputingCAI/safety-for-whom
- Record: https://forck.live/items/9274-safety-for-whom-refusing-the-right-subset-of-a-topic-not-the-whole-topic
- Subject: Platform
- Models affected: Qwen3-8B

The paper introduces boundary-aware self-distillation for controlled LLM safety refusal, focusing on refusing only a harmful subset of a topic rather than the entire topic. Using political persuasion as a testbed, the authors identify weaknesses in standard self-generation safety tuning, including coverage gaps, downside reactions, and inadequate boundary measurement. Experiments on Qwen3-8B show that while political refusal rates improve from 9.47% to 84.75%, over-refusal on safe prompts rises from 2.00% to 74.00%, highlighting the trade-off between safety and over-refusal.

## Evidence

Verbatim from https://huggingface.co/blog/MultiverseComputingCAI/safety-for-whom:

> Our latest paper, Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal , studies this narrower problem directly.

---

Record: https://forck.live/items/9274-safety-for-whom-refusing-the-right-subset-of-a-topic-not-the-whole-topic
Catalogue: https://forck.live/llms.txt
Current issue: https://forck.live/feed.md
