What the AI Text Detection Business Gets Wrong About Its Own Product
OpenAI killed its own AI detector for low accuracy back in 2023. This Business Audit checks what Turnitin, the Authors Guild, and Stanford actually found.
A stranger’s website reads a piece of writing for thirty seconds and returns a percentage, and that percentage now routinely outranks the writer’s own memory of having written the thing themselves. This is the business model behind AI text detection in 2026, and it has grown into a real and profitable industry, licensed to universities, publishers, and hiring platforms, on the strength of a confidence the tools cannot actually back up. The honest audit of AI text detection is not whether AI-generated writing is a real problem. It is. The honest audit is that the businesses selling certainty about it cannot produce that certainty for their own product, and the paper trail confirming this comes largely from the companies themselves.
What the Detector Companies Admit About Their Own Tools
Start with the company that should know the material best. OpenAI built its own AI Text Classifier and launched it on January 31, 2023, with a caveat printed directly into the announcement: the tool correctly flagged only twenty six percent of AI-written text in the company’s own challenge set. By July 20, 2023, OpenAI pulled the tool entirely, citing its low rate of accuracy in the retraction notice appended to that same page. The company that trained the models being detected could not build a reliable detector for them, and said so in writing. Turnitin has a quieter version of the same admission. Its own help center documentation states plainly that it will not display a specific score under twenty percent, because internal testing found a higher rate of false positives in that exact range. Turnitin’s own blog goes further, stating a sentence level false positive rate of roughly four percent, and noting that these misfires cluster disproportionately in the introductions and conclusions of documents, the two sections a writer is most likely to revise by hand. That is not a footnote. That is the company selling the detector describing, in its own published language, where its own results become unreliable enough to hide.
What Independent Testing Actually Found
The Authors Guild ran the test the industry itself never seemed eager to run. In May 2026, the Guild submitted ten of its own articles, every one published years before generative AI existed, into five popular detection tools. Pangram and Originality.ai handled the test well, returning close to zero percent across the board. Grammarly came close behind them, scoring zero percent on eight of the ten articles and only seven and nine percent on the remaining two, low enough that neither would likely trigger concern in practice. ZeroGPT swung wildly, from five percent on one article to seventy six percent on the Guild’s own obituary for novelist Joan Didion, published in 2021, four years before ChatGPT existed. And one tool, Sidekicker.ai, flagged all ten articles as predominantly AI written. Two scored a perfect one hundred percent, including that same Didion obituary. A detector confidently labeled a dead novelist’s memorial as machine output, years after it was written and years before the machines existed to write it.
That result is not an outlier from a small sample. A peer reviewed study by Weber-Wulff and seven co-authors, published in the International Journal for Educational Integrity in 2023, ran fourteen detection tools, including Turnitin itself, through seven hundred fifty six individual tests. Not one tool cleared eighty percent accuracy. Only five crossed seventy. Six of the fourteen produced false positives, thirteen of the fourteen produced false negatives, and accuracy degraded further whenever the text had been paraphrased or run through machine translation. The researchers’ own stated conclusion was blunt: the tools were neither accurate nor reliable, and their errors ran in both directions at once.
The fair counterargument deserves stating plainly rather than skipped past. A narrower 2025 study tested three detectors, GPTZero, ZeroGPT, and a tool called Corrector App, against a thousand academic texts split between pre-ChatGPT human writing and machine generated abstracts, and found accuracy scores between ninety six and one hundred percent under those controlled conditions. That result does not overturn the broader finding here. It sharpens it. These tools can perform well on a narrow, well behaved dataset built specifically for testing them. The moment the input resembles real writing, a paraphrased sentence, a non-native speaker’s phrasing, a stylistically confident human paragraph, the accuracy Weber-Wulff’s team documented is the accuracy that actually governs the disciplinary hearing.
Here is the part that should actually concern a reader, not the part that makes for an easy joke about a robot failing to recognize a dead author. A separate Stanford study, published in the journal Patterns in July 2023, tested seven widely used GPT detectors against two groups of human written essays, one from native English speakers and one from non-native speakers taking the TOEFL exam. The detectors correctly identified the native writers as human almost every time. They misclassified the majority of the non-native writers as AI generated. Senior author James Zou told the press his team’s recommendation was blunt: avoid using these detectors wherever possible, because the consequences fall hardest on students, job applicants, and writers who already face enough scrutiny before a machine adds its own.
Why the Errors Run in This Particular Direction
Why does this keep happening in the same direction, penalizing control rather than carelessness? Because the detectors were never actually reading for authorship. They were reading for resemblance to their own training data, and that training data was scraped from decades of clean, controlled, professionally edited English prose. Emily Dickinson and Friedrich Nietzsche both wrote with the em dash long before any language model existed, and a detector trained on smoothness has no reliable way to distinguish a nineteenth century poet’s discipline from a chatbot’s default output. Agatha Christie died in 1976. Her sentences still get flagged by the same logic people now apply to ChatGPT. The tool cannot tell the difference between a writer who mastered clarity over a lifetime and a machine that was built to imitate exactly that clarity, because at the statistical level these tools actually operate on, that difference may not exist for them to find. Call this the Polish Penalty: the mechanism by which a detector flags controlled, well-edited human prose as artificial, because the model was trained to recognize precisely that style as the signature it is hunting for. The better a person writes, the more they resemble what the machine was told to catch, and no future update quietly fixes that, because it describes the entire architecture working exactly as designed.
What the Business Model Costs, and Who Pays It
None of this means AI generated content is not a real problem worth taking seriously. It is, and publishers, teachers, and readers have every right to want disclosure. But a business built on selling certainty it cannot manufacture is not solving that problem. It is creating a second one, where a wrongly flagged writer has almost no way to prove a negative, and the company holding the percentage carries no liability for being wrong. Run a four percent sentence level error rate across a large enough platform, and the raw count of falsely accused writers stops being a rounding error and starts being a business line nobody wants to name out loud. The money involved is not small, either. Universities pay licensing fees scaled to enrollment, publishers pay for editorial screening tools, and hiring platforms increasingly bundle detection into applicant tracking systems a candidate never sees and never consents to. Every one of those contracts gets renewed on the strength of a headline accuracy figure, often close to ninety eight percent in a company’s own marketing copy, that independent testing under real conditions has repeatedly failed to reproduce.
So what should actually happen here? Any institution using one of these tools to make a consequential decision, expelling a student, killing a book deal, rejecting a job application, should be required to disclose the specific detector’s documented false positive rate alongside the score, in the same sentence, every time. A percentage without its own error margin is not evidence. It is a number wearing the costume of one, and that costume does not belong in front of a disciplinary board, an admissions committee, or a publishing contract.
There is a further, quieter cost that rarely makes it into the coverage of any single detector’s failure. The Authors Guild’s own reporting on this controversy notes that some writers, aware their work might be screened, have begun changing their natural style specifically to avoid sounding like a machine, trimming sentence variety, avoiding the em dash, softening the exact polish that used to be the goal. That is the industry succeeding at something, just not the thing it claims to be selling. A detection business whose side effect is teaching skilled writers to write worse has inverted its own stated purpose, and no accuracy percentage on the label changes that arithmetic.
Demand the error margin. Ask which detector, which version, and which independent study actually backs the figure on the screen. None of that is complicated, and remarkably little of it is currently required anywhere the decision actually gets made. That is the whole audit, and it costs the institution nothing but the discomfort of asking.
Sources for further reading
- OpenAI, “New AI classifier for indicating AI-written text,” original announcement (January 31, 2023), stating a 26% true-positive rate on the company’s own challenge set, with the July 20, 2023 retraction notice appended: “the AI classifier is no longer available due to its low rate of accuracy.”
- Turnitin Help Center, “Why is the AI Writing Detection report score showing as *%?” and Turnitin’s own blog, “Understanding AI writing detection: False positive rates,” confirming a roughly 4% sentence-level false positive rate concentrated near document introductions and conclusions.
- The Authors Guild, “Can AI Detectors Be Trusted? The Authors Guild Put Five of Them to the Test,” May 26, 2026.
- Weber-Wulff, Anohina-Naumeca, Bjelobaba, Foltýnek, Guerrero-Dib, Popoola, Šigut, and Waddington, “Testing of Detection Tools for AI-Generated Text,” International Journal for Educational Integrity, 19, Article 26 (2023) (DOI 10.1007/s40979-023-00146-z); 14 tools, 756 tests, none reached 80% accuracy.
- Liang, Yuksekgonul, Mao, Wu, and Zou, “GPT Detectors Are Biased against Non-Native English Writers,” Patterns, 2023, and ScienceDaily’s July 10, 2023 coverage quoting senior author James Zou of Stanford University.
- A 2025 controlled study of GPTZero, ZeroGPT, and Corrector App against 1,000 academic texts, cited for contrast, showing 96–100% accuracy under narrow, non-adversarial test conditions.
- Rolling Stone, “‘ChatGPT Hyphen’: Are Em Dashes a Giveaway of AI Writing?” (on Emily Dickinson and Friedrich Nietzsche’s historical use of the em dash).
If you found some value in this in depth analysis, you might also want to know Why Your AI Needs a Strict Indian Dad and a Two-Hundred-Dollar Bribe.


One Comment
Comments are closed.