From the source
Lead story
Top stories
Models & availability
Latest
Lead story
Top stories
Models & availability
Latest
From the source
New method trains models to refuse only harmful subsets of topics, not entire categories.
The paper introduces boundary-aware self-distillation for controlled LLM safety refusal, focusing on refusing only a harmful subset of a topic rather than the entire topic.
Using political persuasion as a testbed, the authors identify weaknesses in standard self-generation safety tuning, including coverage gaps, downside reactions, and inadequate boundary measurement.
Experiments on Qwen3-8B show that while political refusal rates improve from 9.47% to 84.75%, over-refusal on safe prompts rises from 2.00% to 74.00%, highlighting the trade-off between safety and over-refusal.
From the source
Our latest paper, Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal , studies this narrower problem directly.
huggingface.co