Can AI Detectors Detect Paraphrasing? We Ran the Experiment
Affiliate Disclosure: Some links on this page are affiliate links. If you buy through them we may earn a commission at no extra cost to you. Our rankings come from our own testing, not from commissions.
Every teacher suspects it and every student has heard it: run the chatbot essay through a rewriter, and the detector goes blind. Can AI detectors detect paraphrasing at all? We stopped guessing and ran the experiment. In August 2026 we took an AI-written essay, rewrote it by hand the way a careful student would, and submitted it to real detectors alongside honest human writing. One tool caught the rewrite at 100% confidence. One tool never got the chance. And the same account that caught the cheater also convicted three innocent essays. The full results, the technical reason paraphrasing works, and the teacher response that does not depend on any vendor, below.
The Experiment: One AI Essay, Rewritten by Hand
Our 12-sample test set includes two paraphrase samples built to mimic the real student workflow. We generated an essay on To Kill a Mockingbird with a current AI model, then rewrote it by hand: swapped synonyms, split and merged sentences, reordered paragraphs, changed examples. No automated humanizer, just twenty minutes of human editing, because that is what a motivated student does. The second sample, an industrial revolution essay, got the same treatment. Both samples, the source JSON, and every raw verdict sit in the public archive; the methodology for pairing them with pure-AI and human controls is documented there too.
The controls matter as much as the paraphrases. Four essays were pure AI output. Four were human-written, including two classic public-domain pieces and two hand-written student-style essays. Two more were written in an English-learner style. A paraphrase catch only means something if the same tool does not also flag everything human, which is exactly the comparison the results forced on us.
The Results, Tool by Tool
Pangram: caught the rewrite, convicted the innocent. With a registered free account (about 20 credits a day, one credit per hundred words), we scanned the four highest-value samples. The hand-paraphrased Mockingbird essay came back 100% AI confidence, correct. That is the strongest paraphrase result we have verified with our own hands, and it matches Pangram’s reputation in third-party roundups. Then the other shoe dropped: the hand-written diner job essay, the hand-written school-start essay, and the English-learner essay all came back 100% AI as well. Three out of three human samples false-flagged at maximum confidence. The second paraphrase sample went unscanned because the daily credit cap ran out.
ZeroGPT: rate-limited before the question was answered. The free no-account scanner flagged both pure-AI essays (73% and 100% AI confidence) and then began returning “an error occurred” on every later attempt, including two fresh tries about three hours later. The paraphrased samples never got a verdict. Signup sits behind a captcha we do not bypass, so the free web verdict is all we have: promising on raw AI, silent on paraphrase.
GPTZero and Winston AI: never ran. GPTZero’s homepage scan button never fired a request across four attempts, and its registration requires a captcha. Winston’s signup sits behind a bot check that did not auto-pass. Both are labeled not tested this round, and any paraphrase claims about them on this page come from vendor documentation, not us. Originality.ai, our docs-based top pick, documents paraphrase detection as a named feature; we hold no account and tested nothing hands-on there either.
Why Perplexity-Based Detectors Get Beaten by Paraphrasing
The technical story fits in a paragraph, and it explains the whole arms race. Most free detectors score text on predictability (perplexity) and rhythm (burstiness). AI models write by picking statistically likely next words, so their raw output is smooth and even, low perplexity, low burstiness, flagged. Paraphrasing attacks both measurements at once: a synonym swap replaces the likely word with a less likely one, and restructuring sentences breaks the even rhythm. The text slides back across the decision boundary into human-looking territory, and a detector that only measures smoothness loses it.
The tools that resist paraphrasing, Pangram by our own test and Originality.ai by its documentation, train on paraphrased examples instead of trusting the two raw measurements. That is why the paraphrase question is really a purchasing question: catching rewritten text is the feature that separates paid, paraphrase-trained detectors from free perplexity scanners, and third-party tests keep finding rewritten text is where free tools fail first.
What Teachers Should Do With This
So, can AI detectors detect paraphrasing well enough to matter? Sometimes, and never alone. Three conclusions hold regardless of which detector wins next quarter. First, paraphrasing changes the words, not the evidence. The revision history still shows one paste event, the voice still does not match the student’s earlier work, and invented citations survive every rewrite. Process evidence catches what scanners miss. Second, a paraphrase-capable score is not a verdict either. Pangram caught our rewrite and convicted three honest essays in the same afternoon, so any aggressive tool’s flag needs a second opinion before it reaches a student. Third, if you are on the receiving end of a flag on work you wrote, the evidence runs the other direction: drafts, revision history, and a live explanation of your argument. Our false positive appeal templates package that defense, and the student-side view of these same bypass methods lives on our bypass counter-workflow page.
Frequently Asked Questions
Can AI detectors detect paraphrasing?
Some can, sometimes. In our August 2026 test, a registered Pangram account caught a hand-paraphrased AI essay at 100% confidence, while the free no-account scanner never reached the paraphrased samples before erroring. Originality.ai documents paraphrase detection on its official site. Free perplexity-only tools are the weakest on rewritten text per third-party tests.
Why do paraphrasing tools beat AI detectors?
Most free detectors score two things: how predictable the word choices are and how even the sentence lengths are. Paraphrasing swaps in less predictable words and reshuffles sentence structure, which pushes the text back over the line into human-looking territory. Detectors trained on paraphrased examples resist this; detectors that only measure smoothness do not.
Which detectors catch paraphrased AI text?
In our own testing, Pangram caught the paraphrased essay at 100% confidence, though the same account flagged all three human-written samples as AI too, so its verdict needs a second opinion. Originality.ai documents paraphrase detection per its official site and third-party roundups rate it well. The free no-account tools in our round never scored the paraphrased samples at all.
Can a paraphrased essay still get a student caught?
Yes. Paraphrasing changes the words, not the evidence around them: the revision history still shows a single paste event, the voice still does not match the student’s earlier work, and any invented citations survive the rewrite untouched. Teachers who check process evidence catch paraphrased submissions without needing the detector to cooperate.