Accuracy & honesty

Honest about what it can and can't do.

A score is a reading, not a ruling. We show the bands plainly, and we tell you where we're wrong — because false accusations are the real harm.

The 0–100 scale, honestly labelled
Reads humanSome AI patternsLeans AIReads AI

Bands, not verdicts: the middle of this scale is genuinely uncertain territory, and we say so. No score here is proof about a person — it's a signal about the text.

Reads human
Some AI patterns
Leans AI
Reads like AI
No detector is 100% accurate — including this one. Formal, technical, or non-native human writing can score high; edited or newer AI can score low. Treat the score as one signal, never as the sole basis for a decision about a person, and never to accuse.
Where it's wrong

The honest failure modes.

01

Formal & non-native writing

Formal, academic, or non-native English writing uses fewer contractions and more even sentences, so it can read high. We ease off those signals by register, but the risk never fully disappears.

02

Edited or newer AI

A light paraphrase or edit removes most surface tells, and the newest human-tuned models can slip past almost entirely. A low score means the text reads human — not that it was.

03

Short text

Under ~300 words most signals abstain or wobble. Give it a few solid paragraphs of writing for the most reliable read.

How to read a score responsibly →

FAQ

Accuracy questions, answered straight.

How accurate are AI detectors, really?

Nobody can honestly give you one number. Published accuracy claims come from each vendor's own test set, and independent studies show every detector — including the biggest names — missing rewritten AI and flagging some human writing. Accuracy varies hugely by text type, length and how much a person edited. That's why this page documents failure modes instead of advertising a percentage.

Why doesn't ScanForAI publish an accuracy percentage?

Because the number would be true only for our test set, and you'd read it as true for your text. A '99% accurate' claim measured on unedited model output says nothing about a lightly-edited essay from a non-native writer — the case that actually matters. We show the evidence per scan instead, so each result argues for itself.

What causes false positives in AI detection?

Formal register, very even sentence lengths, careful grammar, few contractions, template-like structure — the habits of disciplined writing overlap the habits of machine writing. Non-native English writers get hit hardest, on every detector. It's the #1 reason a score should start a conversation, never end one.

What does a mid-range score mean?

Genuine uncertainty — and we label it that way instead of rounding to a verdict. Mid-range text usually mixes signals: some machine-leaning patterns, plenty of human ones. Look at which lines are marked and judge those, not the number.

Which AI detector is the most accurate?

There's no honest king of that hill — rankings flip depending on the test corpus. The strongest signal available to you is agreement: when two detectors with different methods mark the same passages, that's worth far more than either one's score alone. We're built to be a good second opinion: free, and we show our reasons.

AI detection is not 100% accurate. No detector — including ScanForAI — can prove whether text was written by a human or by AI. Every result is an estimate from writing patterns and can be wrong both ways. Never use it as the sole basis for academic, disciplinary, hiring, or other consequential decisions about a person, and never to accuse someone.