Turnitin says its AI detector is highly accurate. Universities have publicly disagreed. Journalists testing it found mistakes. All of these can be true at once — accuracy claims live and die on definitions — so here's the evidence laid out plainly, and what it actually means if your work goes through the detector this semester.
What Turnitin claims
Turnitin has consistently stated its detector is tuned to prioritize avoiding false accusations: a document-level false-positive rate it puts at under one percent for documents scoring above its 20% threshold, which is also why anything below that threshold displays as an asterisk rather than a number. Those are the company's own figures, produced on its own test sets — worth knowing, and worth the grain of salt any vendor benchmark deserves.
What independent testing found
- Mixed documents are the weak spot. Journalists and researchers who tested the detector shortly after launch (including a widely read Washington Post experiment) found it handled purely-human and purely-AI text reasonably well but stumbled on the realistic case — essays that blend human writing with AI assistance — both missing AI text and flagging human sentences near it.
- The bias problem is documented. Peer-reviewed research — "GPT detectors are biased against non-native English writers" — found detectors as a class misclassify a majority of essays by non-native English writers as AI-generated. We cover who gets flagged and why separately.
- Some universities opted out. Most prominently, Vanderbilt University publicly disabled Turnitin's AI detector in 2023, citing the inability to validate its accuracy and the harm of false accusations; a number of other institutions followed or never enabled it. Plenty of others kept it on — which is why your experience depends heavily on where you study.
The math that matters: low rate × huge volume
Here's the part both sides of the argument tend to skip. Suppose the sub-1% false-positive claim is exactly right. Turnitin processes tens of millions of submissions a year — at that volume, "under one percent" still means tens of thousands of wrongly flagged papers, each one a real student having a very bad week. A low error rate and a large number of victims are not contradictory; they're arithmetic. That's the strongest honest case for both positions: the detector is right far more often than it's wrong, and being wrong rarely is still a serious problem at scale.
What this means for you
Three practical conclusions. If you're accused falsely, accuracy evidence is your context, not your defense — your version history and receipts are the defense. If you're submitting AI-assisted work, don't bet on the error rate; detectors miss things, but policies judge what you submitted, not what got caught. And if you simply want to know where your document stands, the detector's accuracy debate is beside the point — running the real check shows you the exact report your marker would see, which is the only "accuracy" that affects your week.