Do AI detectors actually work?
AI writing detectors promise a number: this text is 87 percent likely to be AI-generated. The number is real in the sense that a model produced it. It is not real in the sense that it tells you who wrote the document. The gap between those two things is where the accusations happen.
What the measurements show
Two findings define the field, and they point the same way.
The first is about baseline accuracy. A 2026 study published in the International Journal for Educational Integrity, by Hadra and colleagues at Sultan Qaboos University, tested commercial detectors against four kinds of text: authentic student writing, professional human writing, raw AI output, and hybrid compositions built from roughly half human and half AI content. Overall accuracy landed between 61 and 69 percent. On the hybrid texts, accuracy collapsed toward zero.
That last result is the important one, because hybrid is the normal case. Almost nobody pastes a model’s first draft and submits it. People draft and ask for edits, or write and ask for a rewrite, or generate and then fix. Detectors are tuned for the case that has largely stopped happening.
The second finding is about who gets flagged. In 2023, Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou at Stanford published a study in Patterns that ran seven widely used GPT detectors over TOEFL essays written by non-native English speakers under supervised examination conditions. More than 61 percent of those essays were classified as AI-generated. The same detectors were close to perfect on essays written by US eighth-graders.
The mechanism explains the bias. Most detectors lean on perplexity, a measure of how predictable each next word is. Writing that uses a smaller vocabulary and more conventional sentence patterns scores as low-perplexity, and low perplexity is what these tools read as machine-written. Someone writing carefully in a second language produces exactly that signature. So does a technical writer following a style guide, which is why accuracy on scientific prose runs far below accuracy on humanities prose.
Why watermarking is a different thing
Detection and watermarking get discussed together and work in opposite directions.
A detector looks at finished text and guesses, with no knowledge of how it was made. A watermark is a signal inserted at the moment of generation by the model’s own provider, so verifying it is a lookup rather than an inference. When a watermark is present and intact, it is far stronger evidence than any detector score.
The catch is what absence proves, which is nothing. Older content carries no mark. Not every model watermarks. Provenance metadata comes off with a screenshot, and statistical marks degrade under heavy rewriting. Since August 2, 2026 the EU AI Act has required providers to mark synthetic output in a machine-readable format, and California’s SB 942 imposes parallel duties, so coverage is growing. It will never be complete, and an unmarked document remains unexplained rather than exonerated.
What this means in practice
For schools and universities, the working rule that survives the evidence is that a detector score is a reason to look, never a reason to conclude. Drafting history, version timestamps and a conversation about the work are what actually distinguish authorship. Policies that let a percentage trigger a misconduct process will produce wrongful accusations, and the published false-positive rates make clear whose work gets flagged: international students and anyone writing in a second language.
For publishers and platforms, the calculation is different, because the goal is usually volume management rather than adjudication. Substack added built-in AI detection powered by Pangram in July 2026, and LinkedIn shipped a slop report button that had been clicked more than a million times by August. Neither has to be right about any individual post to be useful at scale, which is the honest case for these tools: filtering, not verdicts.
For anyone being evaluated, the practical defense is boring and effective. Keep drafts. Work in a document with version history. The record of how a piece was written answers the question that a detector only guesses at.
The trend line does not favor detection. Models are getting better at producing varied, unpredictable prose, which is the exact property detectors use to identify human writing. A Pew analysis in August 2026 found signs of AI in about a third of web pages published since ChatGPT launched. As the share of mixed human and machine text rises, the category the detectors handle worst becomes the category that describes most writing.
Related coverage
- Substack added built-in AI detection powered by Pangram, detection deployed as a platform filter.
- Pew finds a third of pages since ChatGPT show signs of AI, the scale of the mixed-authorship problem.
- LinkedIn’s AI slop button has been clicked more than a million times, what readers do when given a flag of their own.
- California’s AI transparency law now requires labels and detectors, detection written into law.
- South Korea indicted a man for wearing AI glasses in an exam, where the enforcement problem is moving.
- How AI watermarks work, and why they can be removed, the provenance side of the same question.
Quick answers
How accurate are AI detectors?
A 2026 study in the International Journal for Educational Integrity by Hadra and colleagues at Sultan Qaboos University measured commercial detectors at 61 to 69 percent overall accuracy against a mix of authentic student writing, professional human text, raw AI output and hybrid compositions. On hybrid text, where a human and a model both contributed, accuracy fell toward zero.
Do AI detectors discriminate against non-native English speakers?
The evidence says yes. A Stanford study by Weixin Liang and colleagues, published in Patterns in 2023, tested seven widely used detectors on TOEFL essays written under supervised exam conditions and found more than 61 percent were misclassified as AI-generated, while the same detectors scored near-perfect accuracy on essays by US eighth-graders. The authors attribute it to lower perplexity in non-native writing, meaning simpler and more predictable word choice, which is exactly what detectors read as machine output.
Can you beat an AI detector by editing the text?
Yes, and it does not take much. Research consistently finds that synonym substitution and light restructuring cut detection rates substantially, and that detectors are built around raw, unedited model output. Mixed human and AI text, the normal case in real use, is where they perform worst.
Can a school punish a student based on an AI detector score?
A detector returns a probability, not proof of authorship, and the published error rates are far too high to treat a score as evidence on its own. Institutions that use these tools responsibly treat a flag as a prompt to look at drafting history, version control and a conversation with the student, not as a finding.