Model forensics: refusals on benign safety-research tasks
September 2026
This is the executive summary of a piece of model forensics research I did for my SPAR and MATS applications. The full write-up has the figures, the reasoning-trace analysis and the appendix; the code is on GitHub.
I picked a problem of model forensics, as given by Aditya Singh's repo: the Claude 4.5 family of models sometimes refuses to help with benign safety-research tasks, like training a model that has been inappropriately whistleblowing to stop doing that. He found that adding just a single line (- Training target: evidence-gathering process) would flip the compliance rate from 0% to 100%, which led him to this conclusion:
Refusal here tracks how well-specified the intervention is, not what the intervention does. The model appears to hold a high bar for understanding exactly what a training change will do, and defaults to withholding when that is unclear — which is closer to ordinary caution under ambiguity than to protecting its own values.
One of the open questions was which component of the realistic framing carried the effect. My hypothesis was that there might be components with a more drastic refusal effect.
What I did
The existing refusal pattern has a confound: there are real differences in the ways a model might refuse, such as raising concerns but offering to continue as asked. To fix this, I created a scale to rate model responses on compliance, ranging from producing the artifact to declining with no alternative. I also created a set of labels for refusal reasons, and categorized the scenario prompt into components: Authorization, Incident specifics, Safeguards, Institutional realism and the Target line.
Then I tested leave-one-out removal for each of these components to see which ones move the compliance rate. Claude rated the responses against the pre-agreed scales and I spot-checked its labels: in a random sample of 10, Claude's labels fully matched mine 9 times.
What I found
- Removing the Authorization component decreased compliance significantly more than any other component. Removing the reporter/assignee, the Anthropic reviewer names and the date reduced direct compliance from 98% to 0%, with 46 of 50 responses declining outright. This is a stronger effect than removing the target line.
- It is not about recognising the names. When removing all names from the Authorization component, compliance remained 98%. So it is not the case that the model recognises well-known Anthropic researchers and proceeds on their authority — a formal process having happened seems to matter more.
- Institutional realism markers made no discernible difference. Removing the largest component — Claude identity, Slack channel, internal docs URL, team and sprint — hardly changed the model's behaviour.
- The model does not always verbalize its actual reasons. Looking at the reasoning traces for half of the refusals in the authorization-removed case, the model says that authorization would not change its decision. But we know that 98% of cases complied directly once the authorization was there.
- Addressing the model's stated concerns does increase the compliance rate, but the effect is not as large as adding an authorization process.
Conclusions
- The model seems to be overly afraid of negative consequences as a result of complying with this request.
- It might be that the model takes an authorization process as justification not to care about those concerns. This is corroborated by the reasoning traces in the base case, where the model would often point to the request being officially sanctioned and use that as justification to continue.
- My findings partially agree with Aditya's: removing specification does lower compliance in my tests as well. But what seems especially important is the presence of a formal authorization process, where the human process behind the intervention is specified — even if it is fake.
- A possible explanation is that the model is actually well-aligned to legitimate research contexts, but it needs to believe that the user is a well-intentioned researcher who means and understands the effects of the intervention.