An AI detector score can indicate that text resembles patterns in a detector's AI examples. It cannot, by itself, prove who wrote the text or how it was produced. A false positive occurs when human-authored writing is labeled as AI-generated.
Why false positives are not just a software bug
Human writing is not one uniform category. Students differ in language background, subject, vocabulary, editing style, disability accommodations, and experience. AI systems are also trained on large amounts of human writing, so the two groups of text overlap.
A 2026 paper frames this as a structural detection problem. In a one-shot, text-only review, an assessor usually does not know the individual student's full writing distribution. The paper argues that any detector with useful power must trade some missed AI text against false accusations where human and machine distributions overlap. Better engineering may change performance, but it cannot make overlap disappear.
What current studies tell us
In a 2025 evaluation using 28 AI-generated and 50 human-written essays of different lengths, GPTZero generally identified purely generated samples but produced fluctuating results for human essays and some false positives. The authors concluded that educators should use caution when relying only on AI detection.
These papers do not establish one universal “false positive rate” for every detector. Accuracy changes with the product version, threshold, model, prompt, sample length, language, subject, and test dataset. A percentage quoted from one study should not be presented as the performance of every tool in every classroom.
Score, evidence, and proof are different
| Item | What it can tell you | What it cannot establish alone |
|---|---|---|
| Detector score | How strongly a passage matches the product's classification patterns | The author, tool used, or whether a policy was violated |
| Version history | How a document changed over time | Every off-platform drafting or assistance step |
| Drafts and notes | Development of ideas, research, and revision | Authorship without considering authenticity and context |
| Student discussion | Understanding of claims, evidence, and writing choices | A complete process record on its own |
A better review process
- Use the detector as a prompt for review. Do not convert the displayed probability into a misconduct probability.
- Inspect the full context. Look at the prompt, allowed tools, passage length, citations, and prior work.
- Ask for process evidence. Review drafts, version history, notes, sources, and feedback.
- Give the student a chance to explain. Ask about the argument and revision choices without assuming guilt.
- Follow a documented policy. Record evidence, decisions, deadlines, and appeal rights.
The BSSS teacher guide provides one example of this evidence-based approach: it discusses drafts, validation tasks, questioning, fair hearing, and evidence of authorship. Institutions should follow their own policies, but transparent multi-source review is more defensible than a detector-only decision.
How students can reduce future uncertainty
Draft in a system with version history, keep outlines and source notes, record permitted assistance, and preserve feedback. These habits support academic work even when no detector is involved. They also provide much stronger evidence than repeatedly editing prose until a third-party score changes.
Students dealing with a current result can start with why an essay gets flagged, then use the false accusation response checklist. For a mechanical edit that does not use generative AI, try the private Non-AI Grammar Checker.
Sources and scope
- Garland, “AI Detectors Fail Diverse Student Populations” (submitted March 2026).
- Dik et al., “Assessing GPTZero's Accuracy in Identifying AI vs. Human-Written Essays” (2025).
- ACT Board of Senior Secondary Studies, Teacher Guide: AI and Academic Integrity.
The two research items above are arXiv manuscripts. They are useful evidence, but they should be read with their methods and limitations rather than treated as a universal benchmark for every detector.