# How Accurate Is ScanForAI? — AI Detector Accuracy, Honestly

What ScanForAI's score can and can't tell you: the bands, the known false positives, and why a score is a signal — never a verdict.

- Source: https://scanforai.com/accuracy
- ScanForAI — AI-writing detection. A score is a signal, not a verdict.

---

# Honest about what it can and can't do.

A score is a reading, not a ruling. We show the bands plainly, and we tell you where we're wrong — because false accusations are the real harm.

Bands, not verdicts: the middle of this scale is genuinely uncertain territory, and we say so. No score here is proof about a person — it's a signal about the text.

## The honest failure modes.

### Formal & non-native writing

Formal, academic, or non-native English writing uses fewer contractions and more even sentences, so it can read high. We ease off those signals by register, but the risk never fully disappears.

### Edited or newer AI

A light paraphrase or edit removes most surface tells, and the newest human-tuned models can slip past almost entirely. A low score means the text *reads* human — not that it was.

### Short text

Under ~300 words most signals abstain or wobble. Give it a few solid paragraphs of writing for the most reliable read.

[How to read a score responsibly →](https://scanforai.com/docs/interpreting-scores)

## Accuracy questions, answered straight.

Nobody can honestly give you one number. Published accuracy claims come from each vendor's own test set, and independent studies show every detector — including the biggest names — missing rewritten AI and flagging some human writing. Accuracy varies hugely by text type, length and how much a person edited. That's why this page documents failure modes instead of advertising a percentage.

Because the number would be true only for our test set, and you'd read it as true for your text. A '99% accurate' claim measured on unedited model output says nothing about a lightly-edited essay from a non-native writer — the case that actually matters. We show the evidence per scan instead, so each result argues for itself.

Formal register, very even sentence lengths, careful grammar, few contractions, template-like structure — the habits of disciplined writing overlap the habits of machine writing. Non-native English writers get hit hardest, on every detector. It's the #1 reason a score should start a conversation, never end one.

Genuine uncertainty — and we label it that way instead of rounding to a verdict. Mid-range text usually mixes signals: some machine-leaning patterns, plenty of human ones. Read the score and the marked lines together — the score is the more reliable half, and the marks show you where the patterns cluster so you can judge them yourself.

There's no honest king of that hill — rankings flip depending on the test corpus. The strongest signal available to you is agreement: when two detectors with different methods mark the same passages, that's worth far more than either one's score alone. We're built to be a good second opinion: free, and we show our reasons.

