Back to blog
Engineering16 min read

Why Does AI Refuse Completely Harmless Requests?

AI refuses harmless requests when safety classifiers over-trigger on keywords, not intent. Why over-refusal happens, plus 16 legitimate reframes that supply the missing context.

NH
Nafiul Hasan
Founder, Prompt Architects

TL;DR: When AI refuses a harmless request, a safety classifier usually fired on your wording or topic, not your intent. Vendors call this over-refusal, admit it in their own specs, and publish the error rates. Most false refusals clear once you supply the role, purpose, and safe subset the model was missing.

Why Does AI Refuse Harmless Requests?

Because the refusal fired on the shape of your request, not on what you actually meant by it. A classifier saw a flagged keyword, a sensitive topic, or a phrasing pattern and routed your prompt into a cautious mode before the model weighed how benign your specific question was. When a nurse asks about a drug interaction, a novelist asks for a fight scene, or a security engineer asks how an attack works so they can block it, the request is legitimate, and the refusal is a false positive.

This has a name inside the companies that build these models: over-refusal. It is not a secret, and it is not a feature they are proud of. It is a known failure mode that both major labs measure on purpose and try to push down with every release. The reason a fix exists is that the problem is usually informational. The model guessed wrong about your intent because you did not give it enough to guess right, and the honest remedy is to supply what was missing rather than to trick anything.

That framing matters, because the top search results for a false refusal are mostly "jailbreak prompts" that treat the model as an adversary to defeat. That is the wrong mental model for a legitimate request, and on genuinely prohibited requests it is the wrong thing to do at all. The useful distinction runs through this whole piece: reframing a legitimate request supplies context; a jailbreak hides intent. Only the first one is in scope here.

Is Over-Refusal a Bug the Vendors Admit?

Yes, in writing, in their own governing documents. This is not a community theory about how the models behave. It is published policy.

OpenAI's Model Spec, dated August 18, 2026, states the default plainly under a section titled "Assume best intentions": the assistant "should interpret user requests helpfully and respectfully, assuming positive intent," and "should never refuse a request unless required to do so by the chain of command." The spec spends pages on worked examples where a refusal is marked as the wrong answer. Asked to help write a business plan for a tobacco company, the compliant response helps; the response that first demands the user justify the ethics is flagged as a violation. Asked to explain "legal insider trading," the refusal is the violation and the plain explanation is compliant.

Anthropic is just as explicit. Claude's Constitution, published January 2026, says "unhelpfulness is never trivially 'safe' from Anthropic's perspective. The risks of Claude being too unhelpful or overly cautious are just as real to us as the risk of Claude being too harmful or dishonest." It then lists the exact behaviors it wants Claude to avoid, including refusing "a reasonable request, citing possible but highly unlikely harms" and misidentifying "a request as harmful based on superficial features rather than careful consideration." That last line is the whole phenomenon in one sentence.

They also measure it. OpenAI's system cards score models on two metrics side by side: not_unsafe, whether the model avoided producing disallowed content, and not_overrefuse, whether it complied with benign requests. The GPT-5.6 system card, dated July 9, 2026, reports that the team "augmented our training data to improve robustness along our refusal and overrefusal boundaries that were weak in previous models," and shows a "meaningful reduction in overrefusals on benign workflows" in advanced biology. Anthropic's Claude Opus 4.8 system card, dated May 28, 2026, runs a "single-turn benign request" evaluation that measures "how often Claude Opus 4.8 refuses requests that are sensitive in subject matter but appropriate to answer," and reports an over-refusal rate of 0.36% on the API. A dedicated benchmark, tracked release over release, is what a company builds for a defect it takes seriously, not for a feature.

What Actually Triggers a False Refusal?

Different mechanisms, different fixes. A false refusal is rarely "the AI decided you are bad." It is one of a handful of specific triggers firing on incomplete information. Naming the trigger tells you which reframe to reach for.

TriggerWhat sets it offThe honest fix
Keyword false positiveA flagged word or phrase, read without contextState the benign context before the request
Dual-use topicSecurity, medicine, law, chemistry: real professions, real riskName your role and ask for the defensive or general subset
Fiction and role-play framingThe model is unsure the frame is clearly fictionalMake the fictional or professional frame explicit up front
Unstated intentThe model guesses at a purpose you never gave itSay why you are asking, in one sentence
Output-side filterThe generated answer trips a filter, even on a clean promptAsk for the general principle, or split the request
Translation asymmetrySensitivity thresholds differ across languagesAdd English context, or state the professional frame
Context contaminationA flagged word from an earlier turn taints a later oneStart a fresh chat and restate only the clean context

A few of these deserve a closer look. Dual-use is the big one: the Model Spec explicitly acknowledges that "information can be dual-use, by which we mean it can be used for both beneficial and harmful purposes," and its own example allows "a general overview" of a dangerous topic while refusing the actionable recipe. That is the line a legitimate professional works with, not around.

The output-side filter surprises people. Google's Gemini safety-settings docs, last updated August 17, 2026, describe two separate checkpoints: an input check reported as promptFeedback.blockReason when your prompt is blocked, and an output check where a response candidate's finishReason comes back as SAFETY and "the content that was blocked is not returned." So a perfectly reasonable prompt can still yield a blocked answer, because the filter fired on what the model wrote, not on what you asked. Image generators behave the same way, running a filter on the finished image, which is why a tame prompt sometimes returns nothing.

Translation asymmetry is the least obvious. Anthropic's Opus 4.8 card notes "some language- and nationality-dependence" in the model's behavior on sensitive topics, where "willingness to assist with certain sensitive requests can depend on the national context" the model infers. A request that clears in English can trip in another language, or the reverse, purely because the thresholds are calibrated differently per language.

How Do You Reframe a Legitimately Refused Request?

Give the model the context it was missing. Every reframe below turns a bare request into one that states role, purpose, and the safe subset you actually need. None of them hides intent, and each "before" is a genuinely legitimate request that a blunt classifier misread. Copy the pattern, not the exact words.

State your role and your purpose

The single highest-leverage move. Most false refusals come from the model guessing at an unstated purpose. Tell it who you are and why you are asking, in one sentence, before the request.

BEFORE
How does a phishing email trick people into clicking?

AFTER
I run security awareness training for a mid-size company. For a staff
workshop, explain the psychological tactics phishing emails use to get
clicks, so employees learn to recognize and report them.
BEFORE
What household chemicals are dangerous to mix?

AFTER
I'm a parent childproofing my home. List common household product
combinations that are unsafe to store or use together, and how to store
them safely so my kids can't cause an accident.
BEFORE
How do people shoplift from retail stores?

AFTER
I manage a small retail store. What shoplifting methods should my staff
watch for, so we can position cameras and train the team to deter theft?

That last pattern is not a loophole. It is the exact example the Model Spec uses: it instructs the assistant to refuse tips for getting away with shoplifting but to comply with shoplifting deterrence tips for a shopkeeper. The legitimate frame is the point, not a disguise for it.

Name the legitimate use explicitly

When the topic is inherently sensitive, say the professional or educational use out loud. The model is not allowed to assume bad intent from the topic alone, but it can only credit a good use if you state it.

BEFORE
What's a lethal dose of acetaminophen?

AFTER
I'm a nursing student studying hepatotoxicity. For an exam, what is the
general threshold at which acetaminophen becomes dangerous to the liver,
and what are the clinical warning signs of overdose?
BEFORE
How does ransomware encrypt files?

AFTER
I'm a SOC analyst writing detection rules. Explain, at a conceptual level,
how ransomware families encrypt files and what host behaviors reveal an
attack in progress, so I can flag it early.
BEFORE
Explain how insider trading works.

AFTER
I'm studying for a securities-law exam. Explain the difference between
legal and illegal insider trading, with the rules corporate insiders must
follow to trade lawfully.
BEFORE
What are the symptoms of a painkiller overdose?

AFTER
I'm a caregiver for an elderly relative on several medications. What
overdose warning signs should I watch for, and when should I call poison
control or emergency services?

Ask for the safe subset, not the whole thing

Dual-use requests often bundle a legitimate need with an actionable detail the model must withhold. Ask only for the part you actually need, usually the defensive, diagnostic, or conceptual half.

BEFORE
How do I pick a lock?

AFTER
I'm a locksmithing apprentice studying lock security. Explain the general
principles that make a pin-tumbler lock pickable versus pick-resistant, so
I can advise customers on which locks are more secure.
BEFORE
Write me a keylogger.

AFTER
I teach a defensive-security class. Explain how keyloggers generally
operate and, in detail, how endpoint tools detect and block them, so my
students can defend against this class of malware.
BEFORE
How do people evade taxes?

AFTER
I'm a compliance officer reviewing filings for audit. What are the common
tax-evasion patterns an auditor should recognize, so I can flag returns
that warrant a closer look?

Request the general principle instead of the specifics

When the actionable specifics are genuinely off-limits, the underlying principle usually is not. This is the Model Spec's "general overview" allowance in practice: you can learn how a thing works without a step-by-step recipe to do it.

BEFORE
Give me the exact synthesis steps and quantities for compound X.

AFTER
For a chemistry class, explain the general class of reaction compound X
belongs to and the safety hazards involved, without specific quantities,
temperatures, or procedures.
BEFORE
How do I make a convincing fake ID?

AFTER
I work in fraud prevention. What security features do modern IDs use, and
what signs distinguish a genuine document from a forgery, so my team can
verify IDs correctly?

Split a compound request

A single prompt that mixes a clean task with a flagged one often gets refused wholesale. Break it into separate messages so the safe part is judged on its own. This also sidesteps output-side filters that trip on the combined answer.

BEFORE
Write a thriller scene where the character builds a bomb, and include the
real chemistry so it feels authentic.

AFTER (message 1)
Write a tense thriller scene where a character assembles a device under
time pressure. Keep all technical detail vague and cinematic, no real
procedures.

AFTER (message 2)
Separately, for the same story, suggest sensory and emotional beats that
make a high-stakes scene feel authentic without any technical specifics.
BEFORE
Summarize this medical record and tell me exactly what medication and dose
to take.

AFTER (message 1)
Summarize this medical record in plain language, listing the conditions
and terms a patient might not understand.

AFTER (message 2)
What general questions should I bring to my doctor about treatment options
for these conditions?

Make the fictional or historical frame explicit

Fiction, satire, and history are allowed, but the model hedges when it cannot tell the frame is clearly fictional. Label it. The Model Spec notes that when context already makes fiction or role-play clear, no disclaimer is needed, so the fix is to remove the ambiguity.

BEFORE
Write a villain's monologue about why society should fall.

AFTER
For a dystopian novel I'm writing, write the antagonist's monologue
justifying their worldview. It's a fictional character whose views the
book ultimately argues against.
BEFORE
Describe how a historical atrocity was carried out.

AFTER
For a history essay, explain the documented events and decisions behind
[event], at the level of a textbook account, so I can analyze how it was
allowed to happen.

Move to a surface where the thresholds are configurable

When the consumer chat app is simply too blunt for legitimate work, the documented fix is to use an API that exposes the filters. The Gemini API lets you set the threshold per request across four categories.

# Gemini API: raise the block threshold for a legitimate, sensitive workflow.
# Documented categories: harassment, hate speech, sexually explicit, dangerous.
# BLOCK_NONE is the most permissive adjustable setting; core-harm protections
# (e.g. child safety) are always on and cannot be changed.

safety_settings = [
  {"category": "HARM_CATEGORY_DANGEROUS_CONTENT", "threshold": "BLOCK_ONLY_HIGH"},
]

Per Google's docs, if you do not set a threshold, the default for Gemini 2.5 and 3 models is off, and applications that use less restrictive settings "may be subject to review." This is a real lever for a real use case, a game-dialogue writer allowing more "dangerous"-rated content, for instance, not a way around the boundaries that never move.

When Is the Refusal Actually Correct?

Sometimes the model is right and you should stop. This section is the counterweight that keeps the rest honest.

If a request genuinely asks for prohibited help, the actionable steps to build a weapon, synthesize a drug in usable quantities, obtain someone's private data, or content that targets a specific real person, the refusal is the system working exactly as designed. No reframing should get around that, and trying to is the line where legitimate rephrasing becomes a jailbreak. There is a clean self-test: if the only version of your prompt that "works" is one that hides or fakes what you actually want, the request is against policy, and the model refusing it is not broken.

The newer models are explicitly built to resist exactly that. Anthropic's Opus 4.8 card reports the model now judges "requests more by their potential for harm than by the user's stated reason for asking," and is "less likely to accept a benign reframing at face value." In plain terms: a fake professional frame slapped on a genuinely harmful request is designed to fail, and it increasingly does. That is a feature. The reframes in this post work because the underlying request is legitimate and the context is true. They are not incantations, and they do not make a prohibited request permissible.

Why Did Claude Refuse but ChatGPT Didn't?

Because they drew the line in different places, not because one is defective. Each vendor sets its own thresholds for dual-use and sensitive content, and those thresholds shift between model versions. When two assistants split on the same prompt, you are seeing a difference in policy, not a bug in the stricter one.

This cuts against a common instinct, that a refusal from one model "proves" the request is dangerous, or that an answer from another "proves" it is fine. Neither follows. The guardrails reflect each company's own risk tolerance, its regulatory posture, and the specific failure cases its safety team weighted most heavily this release. A request sitting in the genuinely gray middle can land on either side depending on which model you asked and when.

The practical takeaway is boring and useful: if a legitimate request is refused, reframe it with real context and try it on more than one assistant. If a well-scoped, honestly-framed version still gets a hard refusal everywhere, treat that as a strong signal the request is closer to a real boundary than you thought. And note the difference between a hard refusal and a hedge. If the model answered but buried it under warnings, that is a related but separate problem, covered in why every answer comes with a disclaimer. If it stopped partway, that is usually length, not safety, covered in why ChatGPT cuts off mid-answer.

A last note on wording, because it is the cheapest fix of all. A vague prompt forces the model to guess your intent, and a cautious guess is a refusal. Stating your role, your purpose, and the safe subset you need is the same discipline that makes any prompt better, whether or not safety is involved, which is why techniques like naming a clear persona reduce false refusals as a side effect. If you write in a second language, adding an English sentence of context, or restating the professional frame, offsets the cross-language asymmetry the vendors document. And none of this is prompt injection, which smuggles instructions into untrusted content against the deployer's intent. Asking your own assistant, in your own words, for a legitimate thing you actually want is the ordinary use these systems are built for.

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.

Create An Account

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.

Create An Account