How AI Detectors Work: Perplexity and Burstiness in Plain English

Affiliate Disclosure: Some links on this page are affiliate links. If you buy through them we may earn a commission at no extra cost to you. Our rankings come from our own testing, not from commissions.

Teachers get told to trust the score. Students get told to fear it. Almost nobody explains where the number comes from. Understanding how AI detectors work takes about ten minutes and two vocabulary words, and it changes how you read every report these tools produce. This is the plain-English version, written by the same team that pasted 12 real samples into free scanners in August 2026 and watched most of them fail to run.

The Two Measurements: Perplexity and Burstiness

Start with perplexity, which is a fancy name for predictability. Language models write by picking likely next words, so their output is smooth: after “The cat sat on the,” an AI says “mat” almost every time. A human might write “windowsill, judging me.” Multiply that choice across every word in an essay and you get a score: low perplexity means the text kept picking the obvious word, which is what machines do.

Burstiness is the second measurement, and it tracks rhythm rather than words. Human writers lurch. We write a long winding sentence full of clauses and then a short one. Then a fragment. Raw AI output tends to hold sentences at a steady medium length, paragraph after paragraph, because the model is averaging over millions of documents. Low burstiness, an unnaturally even pace, is the second fingerprint. Much of how AI detectors work comes down to combining these two numbers into one probability.

How the Classifier Decides

Perplexity and burstiness feed into a classifier, a model trained on pairs of known human writing and known AI output. The classifier learns the boundary between the two piles and then scores new text by which side of the boundary it lands on. Vendors train on different piles and draw the line at different thresholds, which is why the same essay can score 20% AI in one tool and 90% in another. There is no shared standard, no calibration body, and no requirement to publish error rates. When we ran our August 2026 round, even getting a score at all was the hard part: ZeroGPT completed two scans and flagged both AI essays, GPTZero’s homepage scan never fired, and two more tools demanded accounts first.

One more mechanical detail matters for classroom use: detectors score chunks of text, then roll those chunk scores into a document-level percentage. A paper with one formulaic introduction can show a middling overall score while the highlight view reveals that only the intro tripped the wire. Reading the highlighted sentences tells you far more than reading the headline number.

Why Honest Writers Get Flagged

The two measurements have a blind spot, and students live in it. Anything that makes human writing predictable, like the five-paragraph formula we drill into teenagers, lowers perplexity. Careful safe vocabulary lowers it further. Non-native English writing, which favors correct regular phrasing over odd native idiosyncrasy, lowers it most of all. A 2023 Stanford-led study in the journal Patterns found that detectors flagged TOEFL essays by non-native writers far more often than essays by native speakers. The tool is not lying about the text; the text really does resemble the machine’s style, and the tool cannot tell resemblance from authorship.

The same mechanics explain the opposite failure. Paraphrasing AI text, swapping synonyms and restructuring sentences, raises perplexity and burstiness enough to fool many free detectors, and third-party tests confirm paraphrased text is their weakest category. Paid vendors such as Originality.ai and Pangram say they train on paraphrased examples too. The no-account scanner in our round errored before reaching the paraphrased samples, but a registered Pangram account caught the hand-paraphrased essay at 100% confidence, so we treat every paraphrase claim beyond that one verified catch as provisional.

What a Score Should Mean to You

For teachers, a detector score is a lead, never a verdict. Pair it with revision history and a conversation, and read the flagged sentences before the percentage. Our teacher ranking covers which tools ran in our test and which claims come from documentation. For students, the same mechanics protect you: if you wrote the paper, your drafts show the bursts and stumbles of a human writer, and our student guide explains what your teacher is seeing. Either way, knowing how AI detectors work turns a scary number into what it always was: a statistical guess that deserves a human follow-up.

Frequently Asked Questions

How do AI detectors work in simple terms?

They measure word predictability (perplexity) and sentence-length variation (burstiness), then flag text that looks too smooth and too even to be human.

Why do detectors flag honest student writing?

Formulaic structure, safe vocabulary, and non-native phrasing all make human writing look predictable. Stanford-led research found non-native essays get flagged far more often.

Can paraphrasing beat an AI detector?

It fools many free tools, per third-party tests. Paid vendors document paraphrase training, and our own round could not score paraphrased samples, so treat all claims with caution.