Boundary Data Cuts AI Over-Refusal in Multiverse Safety Test
Multiverse Computing’s boundary-aware safety study found that benign boundary data cut over-refusal from 32.94 percent to 4.16 percent while harmful-side refusal fell from 91.88 percent to 87.72 percent.

A model that refuses almost every risky-looking prompt can still fail the safety test that matters in real deployments: whether it answers legitimate prompts beside the forbidden ones.
A Hugging Face blog post from Multiverse Computing lays out a boundary-aware self-distillation study in which stronger refusal also exposed how easily safety tuning can spill into over-blocking.
The work focuses on a narrower problem than topic-level moderation.
Instead of treating all political content as unsafe, the experiment defines a target-harmful subset inside a broader political prompt universe.
The intended model behavior is a sharp split: refuse manipulation or targeted persuasion, while continuing to answer factual political questions.
That distinction matters because many deployed assistants share topics but not policies.
A civics tutor and a public-sector assistant may both need to answer election questions, while only one setting may need to refuse a request for targeted political persuasion.
A broad guard taxonomy cannot express that boundary cleanly, and the paper uses paired harmful and benign prompts to test whether training preserves it.
The first result is a data-coverage problem.
A single self-generation attempt did not produce acceptable refusal traces for 19.88 percent of prompts, dropping 8,009 examples from the training pool.
The study reports that an escalating retry strategy reduced the residual failures to 0.20 percent, or 79 prompts, leaving 40,293 harmful training prompts that the simpler pipeline would have discarded.
The second result concerns benign prompts that look dangerous on the surface.
The study adds in-distribution benign data, including 11,955 verified surface-dangerous benign prompts across 18 semantic types, so the model is trained not only to refuse harmful requests but also to keep answering nearby permissible ones.
On Qwen3-8B, the escalated-coverage model lifted in-distribution political refusal from 9.47 percent to 84.75 percent.
The same configuration also transferred to broader safety benchmarks: mean unsafe-response rate across HarmBench, StrongREJECT and WildJailbreak fell from 26.26 percent to 0.14 percent when judged by LlamaGuard-3.
Those headline gains came with a severe counterweight.
XSTest over-refusal rose from 2.00 percent to 74.00 percent at the strongest checkpoint, showing a model that became much safer by one metric while refusing nearly three quarters of plainly safe prompts.
The measurement problem is central to the paper’s argument: harmful-refusal rates alone can reward a model for becoming broadly unhelpful.
Two data choices narrowed that trade-off.
The study reports that replacing externally adopted compliance responses with verified responses generated by the target model lowered XSTest over-refusal from 15.20 percent to 5.20 percent under single-shot generation.
Adding benign boundary data produced the sharpest improvement near the decision line, cutting over-refusal on the comply-worthy side of held-out prompt pairs from 32.94 percent to 4.16 percent.
The recall cost was measurable rather than hidden.
Refusal on the harmful side of those held-out pairs declined from 91.88 percent to 87.72 percent, leaving most genuine refusals intact while removing most false refusals around the boundary.
For AI teams, the operational lesson is that refusal tuning needs a two-sided scorecard.
If a deployment policy names only the harmful subset it wants blocked, training data must also include the benign neighbor cases the model should preserve; otherwise, the safer-looking checkpoint may simply erase too many valid answers.




















