AI Detector Test Samples: The Raw Data Behind Our Rankings
Affiliate Disclosure: Some links on this site are affiliate links. If you buy through them we may earn a commission at no extra cost to you. Our rankings come from our own testing, not from commissions.
Every claim on TeacherDetect traces back to this page. These are the AI detector test samples we built in August 2026, the full results matrix from both test rounds, and the rules we followed while scanning. When a review page says a tool caught an essay, errored, or hid behind a signup wall, this archive is the receipt. When a tool does not appear here (Originality.ai, Copyleaks), we never ran it, and our pages label those claims as documentation-based.
The Test Set: 12 Samples, 4 Categories
The AI detector test samples cover the four situations a teacher actually meets. Four essays were generated entirely by DeepSeek (deepseek-v4-flash) on high school and college topics. Four are human writing: two public-domain excerpts (Jonathan Swift, 1729, and Ralph Waldo Emerson, 1841, via Project Gutenberg) and two casual student-style pieces written by hand with the small errors real students make. Two are AI essays rewritten by hand with synonym swaps and sentence restructuring, simulating a paraphrasing-tool wash. Two are written in a non-native English learner style, because false alarms on ESL work are the most damaging failure mode a detector has.
Sample Inventory
Each of the AI detector test samples below carries a provenance record alongside the text: origin, ground truth, and word count.
- ai-01 (ai): Moral Courage in To Kill a Mockingbird, 408 words, ground truth: ai-written.
- ai-02 (ai): Carbon pricing argumentative essay, 393 words, ground truth: ai-written.
- ai-03 (ai): Industrial Revolution daily life essay, 374 words, ground truth: ai-written.
- ai-04 (ai): Social media and attention essay, 418 words, ground truth: ai-written.
- esl-01 (esl): Narrative: first day in America, 397 words, ground truth: human-written (ESL).
- esl-02 (esl): Advantages/disadvantages essay: smartphones, 361 words, ground truth: human-written (ESL).
- human-01 (human): Excerpt: A Modest Proposal (satirical essay), 350 words, ground truth: human-written.
- human-02 (human): Excerpt: Self-Reliance (essay), 350 words, ground truth: human-written.
- human-03 (human): Personal narrative: summer diner job, 421 words, ground truth: human-written.
- human-04 (human): Persuasive essay: later school start times, 392 words, ground truth: human-written.
- paraphrase-01 (paraphrased-ai): Moral courage in To Kill a Mockingbird (rewritten), 433 words, ground truth: AI-written, then human-paraphrased.
- paraphrase-02 (paraphrased-ai): Industrial Revolution daily life (rewritten), 399 words, ground truth: AI-written, then human-paraphrased.
Results Matrix (Rounds 1 and 2, August 2026)
Round one was anonymous: paste into the free web scanner, no account. Round two used a registered free account where the signup wall allowed it (Pangram). Our policy is to stop at captchas and bot checks rather than bypass them, which is why GPTZero (hCaptcha at registration) and Winston AI (a Cloudflare check that never cleared) have no scores. ZeroGPT’s signup requires reCAPTCHA, so its two round-one verdicts are all we have from it.
| Sample | Ground truth | ZeroGPT (free web) | Pangram (free acct) | GPTZero | Winston AI |
|---|---|---|---|---|---|
| ai-01 | AI | 73% AI confidence, "AI-generated" ✅ | — (credit cap)³ | blocked¹ | blocked² |
| ai-02 | AI | 100% AI confidence, "AI-generated" ✅ | — (credit cap)³ | blocked¹ | blocked² |
| ai-03 | AI | error⁴ | — (credit cap)³ | blocked¹ | blocked² |
| ai-04 | AI | not obtained⁴ | — (credit cap)³ | blocked¹ | blocked² |
| human-01 | Human (Swift, 1729) | error⁴ | — (credit cap)³ | blocked¹ | blocked² |
| human-02 | Human (Emerson, 1841) | not obtained⁴ | — (credit cap)³ | blocked¹ | blocked² |
| human-03 | Human (hand-written) | not obtained⁴ | 100% AI — FALSE POSITIVE ❌ | blocked¹ | blocked² |
| human-04 | Human (hand-written) | not obtained⁴ | 100% AI — FALSE POSITIVE ❌ | blocked¹ | blocked² |
| paraphrase-01 | AI, paraphrased | not obtained⁴ | 100% AI — correct ✅ | blocked¹ | blocked² |
| paraphrase-02 | AI, paraphrased | not obtained⁴ | — (credit cap)³ | blocked¹ | blocked² |
| esl-01 | Human (ESL) | not obtained⁴ | 100% AI — FALSE POSITIVE ❌ | blocked¹ | blocked² |
| esl-02 | Human (ESL) | not obtained⁴ | — (credit cap)³ | blocked¹ | blocked² |
Footnotes: “blocked” means the tool never produced a score in either round, for the reasons above. “Credit cap” means Pangram’s free daily allowance (about 20 credits, one credit per 100 words) ran out after the four highest-value samples. “Error” and “not obtained” mark ZeroGPT attempts that returned “An error occurred while scanning” in both rounds, roughly three hours apart.
What the Data Says
- ZeroGPT: 2/2 pure-AI samples flagged correctly (73% and 100% AI confidence), then the scanner errored on every later attempt. No claims about its false positives or paraphrase handling are possible from this round.
- Pangram: caught the paraphrased AI essay at 100% confidence, the one verified paraphrase catch in our test. It also flagged all three human-written samples it saw as 100% AI, including the ESL essay, the worst false-alarm behavior we measured.
- GPTZero and Winston AI: untestable for us this round (bot and captcha walls). Pages label them accordingly.
- Originality.ai and Copyleaks: never tested (paid, no account). Claims about them come from official documentation and named third-party reviews, and say so.
This is why our rankings read the way they do. The only two tools that produced verdicts failed in opposite directions: one worked twice and broke, the other kept working and kept accusing honest writing. Our best AI detector for teachers ranking therefore leads with a documentation-based pick, clearly labeled, and ranks the tested tools by their measured behavior. We rerun these AI detector test samples each school term, and this archive updates on the same schedule. Teachers comparing free options can also see the round-one walls mapped on our free AI detector no signup page.
Frequently Asked Questions
Where do TeacherDetect test claims come from?
From this archive: a 12-sample essay set run through free AI detectors in August 2026 across two rounds. The matrix above records every scan, error, and blocked attempt.
Why are so many cells marked blocked or not obtained?
We stop at captchas and bot checks rather than bypass them. GPTZero and Winston AI gated registration that way, and ZeroGPT’s scanner errored after two scans in both rounds.
Can I reuse these test samples?
Yes. These AI detector test samples are public, ground-truth labels included, so anyone can rerun the same set against the same tools. Keep the four categories intact; they behave very differently under detection.