Fairness · a signal, not a verdict

Why AI detectors flag non-native English

careful ≠ machine

If English is your second language and a detector called your writing “AI,” you were probably not doing anything wrong — you were caught by a known, measurable bias. Here is why it happens, what the research found, and how ScanForAI is built to be fairer about it.

What “ESL” and “non-native” mean

ESL stands for English as a Second Language — anyone who learned English after a first language. “Non-native” means the same thing more broadly. It says nothing about how good the writing is; plenty of non-native writers are more precise and more formal than native speakers.

That precision is exactly the problem. Learned English tends to be careful, even, and formal — and to a machine that measures writing statistically, careful and formal looks a lot like machine-written.

Why detectors over-flag it

Most AI detectors score one core thing: how predictable your text is. Machine writing is smooth and low-surprise because that is how language models generate — always reaching for the likeliest next word. The trouble is that several habits of good non-native English produce the same signature:

None of these mean the writing was generated. They mean the surface statistics overlap — and a detector reading only the surface can drift careful human writing up the same scale it uses for machines:

Reads humanReads machine →

What the research shows

This isn’t a hunch. A 2023 Stanford study ran seven widely-used GPT detectors over essays by non-native and native English speakers. The gap was stark:

61%
of TOEFL essays by non-native writers were wrongly flagged as AI, on average across seven detectors — some flagged up to 98%.
~0%
of essays by native English writers were flagged by those same detectors — the bias lands almost entirely on non-native writing.

Liang et al., “GPT detectors are biased against non-native English writers,” Patterns (Cell Press), 2023.

The same study showed the flip side: asking a model to rewrite the essays in more “literary” language dropped the false-flag rate toward zero. So the tools punish plain, careful English and reward fancier writing — the opposite of what fairness would want.

How ScanForAI is built to be fairer

We can’t make the bias vanish — no detector honestly can — but we designed against it on every layer instead of pretending it isn’t there.

Most detectors

One predictability score. Formality reads as guilt. You get a bare percentage with nothing to check it against.

ScanForAI

Reads your register first and eases formality tells. Shows the exact patterns that fired, line by line. A signal to examine — never a verdict.

Honest caveat: even with all of this, formal and non-native writing still scores higher than casual native writing on every detector, ours included. That’s why we show the reasoning and say, on every result, that it is a signal and not a verdict.

If you’ve been flagged and it’s wrong

Being flagged is not proof, and you have more standing than a single score. Practical steps:

Free, anonymous, and per-line: see exactly what fires on your writing — before anyone else runs it through a worse tool.

Scan your writing free →
FAQ

Common questions.

Why do AI detectors flag non-native English writers?

Learned formal English is careful, even and template-shaped: measured sentence lengths, formal vocabulary, few contractions, predictable phrasing. Those are the same surface statistics machine text shows, so detectors that score “predictability” flag non-native writing as AI far more often than native writing. It’s a bias in what the tools measure, not a judgement about the writer.

What does ESL mean in AI detection?

ESL means English as a Second Language — someone who learned English after a first language. It matters here because learned English tends to be more formal and even than casual native English, and that formality overlaps statistically with machine-written text. So ESL and other non-native writing gets flagged as AI more often, even when it’s entirely human.

A detector said my writing is AI, but English is my second language — what should I do?

Stay calm and gather evidence, not just a denial: your draft history and notes, and a per-line report showing which specific patterns fired. Ask which detector was used and its known false-positive rate for non-native writers. No single score is proof, and most institutions know it. Scanning here gives you a breakdown you can point at in that conversation.

Is ScanForAI fair to non-native English writers?

Fairer by design, and honest about the limits. It detects the register of your text and eases the formality signals that punish careful and non-native writing, it shows the specific patterns behind every score instead of a bare percentage, and it frames the result as a signal rather than a verdict. The risk never fully disappears on any detector — but the evidence is on the page, so a high score starts a conversation instead of an accusation.

Can I prove I wrote something if a detector flags it as AI?

No detector can prove authorship either way, so the burden should never sit on the score alone. Keep your drafts and version history, be able to discuss the content and your sources, and bring a per-line breakdown that shows whether the flagged lines read “machine” or just “careful and formal.” Concrete evidence beats one tool’s number.

Does Grammarly or heavy editing make my writing look AI-generated?

It can. Polishing tools smooth exactly the things detectors measure — sentence rhythm gets more even, wording more standard — so heavily Grammarly-edited human text scores higher than your raw draft. If a clean score matters, scan before and after editing: the per-line view shows which “improvements” pushed it toward machine-typical.

Will translating from my first language make my text look AI-written?

Often, yes. Machine translation produces the same smooth, high-probability phrasing that AI detectors flag, so a DeepL or Google-Translate pass can score “AI” even though the ideas are fully yours. Safer: write directly in English and fix grammar afterwards, or rework the translation in your own voice — then scan to see what still reads machine-made.

Do Turnitin and GPTZero flag international students more often?

Every detector that scores predictability shows this skew, and the research is public: the Stanford study found seven major detectors wrongly flagged 61% of non-native TOEFL essays while flagging native essays near 0%. Individual tools differ and versions change, but if English is your second language, assume elevated false-positive risk on any of them — and keep drafts as evidence.

How do I make my essay less likely to be falsely flagged?

Write like yourself, not like a textbook: vary sentence lengths, keep a few of your natural turns of phrase, and don’t sand every edge off with editing tools. Then scan and read the marked lines — fixing the handful that read machine-typical usually drops the score honestly. Avoid “undetectable AI” tricks; they target a moving standard and look worse when discovered.

Should I tell my professor English is my second language?

If AI detectors are in use, context helps you. Mentioning it before submission — or immediately if flagged — reframes formal, careful writing as what it is: learned English, not machine output. Pair it with your draft history and a per-line report, and you’re bringing evidence rather than only a denial.

AI detection is not 100% accurate. Formal, technical, or non-native human writing can score high; edited AI can score low. Never use a result as the sole basis for a decision about a person.