Back to blog
Engineering10 min read

Why AI Writing Gets Flagged (And What to Do Legitimately)

AI detectors flag text using statistical proxies that misfire on non-native English and formulaic writing. What the classifiers actually measure, and what to do if you're flagged.

NH
Nafiul Hasan
Founder, Prompt Architects

TL;DR: AI writing detectors mostly measure how statistically predictable your word choices are, a signal that misfires on non-native English and plain, formulaic writing. OpenAI's own classifier caught only 26% of AI text while wrongly flagging 9% of human text, and a Stanford study found detectors misclassified non-native writers' essays as AI-generated 61% of the time. If you're flagged, the legitimate response is process evidence and disclosure, not a tool that promises to fool the next scan.

Why do AI detectors flag human writing at all?

Because most of them are not detecting "AI," they're detecting statistical predictability, and plenty of human writing is predictable. The dominant signal is called perplexity: a measure of how surprising each word choice is to a language model. Text that closely matches the kind of phrasing a model would generate on its own reads as "low perplexity" and gets flagged, whether a person or a model actually produced it.

OpenAI's own now-retired classifier is the clearest case study, because the company published its real numbers instead of a marketing claim. On launch, OpenAI reported: "In our evaluations on a “challenge set” of English texts, our classifier correctly identifies 26% of AI-written text (true positives) as “likely AI-written,” while incorrectly labeling human-written text as AI-written 9% of the time (false positives)." That is a tool that missed three out of four AI-written samples and still wrongly accused one in eleven humans.

OpenAI didn't quietly deprecate it either. The same page now carries a banner: "As of July 20, 2023, the AI classifier is no longer available due to its low rate of accuracy." Eighteen months from launch to shutdown, from the company with the most direct financial interest in a working detector.

The company's own published limitations are worth reading in full, because every other vendor's tool shares the same structural weak points even when it doesn't publish the numbers. OpenAI's list includes: the classifier "is very unreliable on short texts (below 1,000 characters)"; it "performs significantly worse in other languages" than English; and it is "poorly calibrated outside of their training data", meaning it can be confidently wrong on anything unlike what it was trained on. None of that is specific to OpenAI's tool. It's a description of what perplexity-based detection can and can't do.

Some current detectors go further than a single document-level score. GPTZero's own FAQ describes flagging at multiple granularities: "GPTZero tells you where AI is used at document, sentence, and word level". Its API returns a dedicated field for sentence-level highlighting on top of that. That sounds like more rigor, but it inherits the same weakness at a smaller scale. A single sentence is a much shorter sample than a full essay, and reliability is exactly what every vendor's own documentation says degrades as a sample gets shorter. Scoring at the sentence level means routinely scoring text in the length range every detector handles worst.

Why does the false-positive problem hit non-native English writers hardest?

Because non-native English writers structurally produce lower-perplexity text, and lower perplexity is exactly what these detectors treat as evidence of AI authorship. A 2023 Stanford study, published in the journal Patterns, evaluated seven widely used GPT detectors against 91 real TOEFL essays written by non-native English speakers and 88 essays written by US eighth graders.

The results were stark. The paper states it plainly: 'The detectors demonstrated near-perfect accuracy for US 8-th grade essays. However, they misclassified over half of the TOEFL essays as "AI-generated" (average false positive rate: 61.22%).' The eighth-grade essays are presumed to be written by native speakers. Eighteen of the ninety-one TOEFL essays were unanimously flagged as AI-written by all seven detectors; only one of the eighth-grade essays saw comparable agreement.

Ninety-one and eighty-eight essays is a modest sample, and the authors say so themselves: the paper's own limitations section states "the sample sizes employed in this research are relatively small." What makes the finding hard to dismiss regardless is the consistency across all seven separately-built detectors and the size of the gap, a 61% average false-positive rate on real human writing against near-zero on the comparison group. A small, tightly controlled study with a large, repeated effect is different evidence than a single striking anecdote, and it's the kind of evidence that should change how much weight you put on any one detector's score.

The study's own explanation is mechanical, not mysterious: the flagged TOEFL essays "had significantly lower perplexity compared to the others". The paper concludes that "GPT detectors may penalize non-native writers with limited linguistic expressions." A writer with a smaller working vocabulary in a second language produces more repetitive, more predictable sentence structures, not because they used AI, but because that's what writing in a non-native language looks like. The detector cannot tell the difference, because it was never measuring authorship in the first place.

The same paper found that editing those essays to broaden the vocabulary cut the average false-positive rate roughly in half. That's worth understanding as a diagnosis, not a recipe: it demonstrates that the score tracks word predictability, not humanity or authorship, which is precisely why no one should treat it as proof of either. A tool that a five-minute edit can flip in either direction was never measuring what it claims to measure.

Does formulaic or professional writing get flagged for the same reason?

Plausibly, yes, by the same mechanism, though it hasn't been measured the way the non-native-writer bias has been. If perplexity is the signal, then any writing style that is naturally repetitive or template-driven, a cover letter, a status report, boilerplate customer correspondence, shares the same statistical shape the detectors already mistake for AI output in non-native writers.

That is an inference from a documented mechanism, not a separately verified statistic, and it should be stated as exactly that. What is documented is that the underlying signal doesn't distinguish "predictable because a human writes formulaically" from "predictable because a model generated it." Anyone whose job requires structured, repetitive prose, a paralegal, a customer support lead, a grant writer working from a template, is operating in the range where this signal is least trustworthy.

What the detector measuresWhat it actually tells youWhy it misfires
Perplexity (word predictability)Statistical unpredictability, not authorshipNon-native and formulaic writing are naturally more predictable, which reads as "AI"
Text lengthConfidence in the score, not accuracyOpenAI called its own classifier "very unreliable on short texts (below 1,000 characters)"
Similarity to training dataHow close the sample is to what the model has seenOpenAI's classifier is "poorly calibrated outside of their training data", which includes non-English text

What should you actually do if your writing gets flagged?

Offer evidence about your process, not an argument about the score. GPTZero's own guidance to educators, the people most likely to be looking at a flagged result, is explicit that a score alone isn't sufficient: "There always exist edge cases with both instances where AI is classified as human, and human is classified as AI." Its documented limitations state plainly that "these results should not be used to punish students."

The same guidance tells you what actually counts as evidence: "Ask the student if they can produce artifacts of their writing process, whether it is drafts, revision histories, or brainstorming notes." A document with an edit history, in Google Docs, in a word processor's version history, or in dated draft files, shows a piece assembled over time rather than pasted in whole. That is a far stronger answer than resubmitting the same text to a different checker and hoping for a better score.

Disclosure is the other legitimate lever, and it only works if you use it before you're asked. If a policy requires you to note where AI helped, whether that's a course syllabus, a publication's submission guidelines, or an employer's content policy, say so up front. Academic style guides have already caught up to this: APA publishes official guidance on citing AI-assisted work, last substantively updated in September 2025, precisely because "I didn't disclose it" is a much worse position than "I disclosed it and here's exactly what the tool did."

In practice, disclosure doesn't need to be dramatic. A one-line note in a methods section, a footnote on a report, or a sentence in a project brief stating that a draft was AI-assisted and then edited and fact-checked by you covers most policies that ask for it at all. The point isn't ceremony. It's having an honest answer ready before a detector score forces the question.

None of this works if a single score is treated as a verdict rather than a prompt to look further. GPTZero's own API documentation says as much about its most granular signal: "The sentence-level classification should not be solely used to indicate that an essay contains AI". If the vendor selling the detector says its own output shouldn't be the sole basis for a decision, that instruction is worth taking at face value, whichever side of the accusation you're on.

Is there a real fix, or just better odds?

Writing in your own voice is the closest thing to a real fix, and it's worth doing for reasons that have nothing to do with any detector. Formulaic, hedge-everything, symmetrically-structured prose reads as machine-written because it's genuinely generic, not because a classifier is clever. Our guide on the specific tells that make writing sound like AI covers the actual habits, reflexive praise, uniform sentence rhythm, safe hedging, and how to prompt your way out of them when AI is helping you draft.

The distinction that matters is intent. Editing a draft so it sounds more like you, with your own examples, your own opinions, and your own sentence rhythm, is good writing practice whether or not any detector exists. Editing a draft specifically to defeat a scanner, without changing anything about whether AI produced the underlying argument, is the thing this article has deliberately not told you how to do.

A message template for contesting a false flag

Use this shape if you wrote something yourself and it was flagged anyway. It leads with evidence, not with an argument about the tool's accuracy.

Subject: Requesting a review of the AI-detection result on [document name]

I'm writing to contest the AI-writing flag on my submission. I wrote this
myself, and I can provide evidence of my process:

- Version/revision history from [Google Docs / Word] showing edits over
  [timeframe], not a single paste.
- Draft files, outlines, or research notes dated before submission.
- I'm glad to discuss the sources or argument directly if that helps.

I understand these tools rely on statistical signals that are documented
to misclassify non-native English writing and plain, direct prose. I'd
appreciate the chance to show my process rather than have the score
stand as the only evidence.
Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 4.8★ on the Chrome Web Store.

Create An Account