How do AI detectors work? The four signals behind the score

Most advice about beating a detector targets vocabulary, which is why it does not work. What moves the score is structural, and so is the fix.

Written by Avoid Content Team · Verified by the seven-stage pipelineSep 18, 20268 min read

ReceiptVerified · 7 stagesSources16Claims12 of 12SEO80AI search84Quality92AI tells0

The word list was never the test

Swap every “delve” for “explore” and the detector score barely moves. The word was never the thing being measured.

A detector holds no blacklist of suspicious vocabulary. It answers a statistical question: whether this passage sits where a language model’s output usually sits, or where human writing usually sits. That question has a real answer, a known accuracy limit, and a failure mode that costs working writers money. All three are worth knowing before you build an editing process on top of a number.

What a detector actually measures

Language models write one token at a time, drawing each from a probability distribution shaped by everything before it. Detection works off the shape of that distribution. According to Eric Mitchell and four Stanford co-authors in DetectGPT, “text sampled from an LLM tends to occupy negative curvature regions of the model’s log probability function.” Machine text sits in a place human text mostly avoids, and you can test for it without training a classifier, a model taught to sort text into two buckets.

Commercial tools wrap that idea in a trained model. GPTZero describes its own product as “a sentence-by-sentence classification model” that scores each sentence for the probability it was written by AI, then rolls those scores up into a document verdict.

Four signals do most of the day-to-day work. Two of them are the detector vocabulary itself: perplexity, a measure of token predictability, and burstiness are the pair Edward Tian named when he built GPTZero. The other two, opener diversity and stock-phrase density, are what a reader notices and what an editor, or a quality gate, can count on a draft:

SignalWhat it measuresHow an unedited AI draft usually scoresWhat an editor changes
BurstinessVariation in sentence length across a passage, as a coefficient of variationLow. Sentences cluster around one comfortable lengthSplits and joins sentences on purpose, so rhythm tracks meaning
Opener diversityShare of sentences that begin with a different word or constructionLow. “The”, “This” and “It” open paragraph after paragraphRewrites openers with a name, a number, a question, a verb-first line
Token predictabilityHow closely each next word matches the model’s expectation, reported as perplexityPredictable. Perplexity sits below the human rangeAdds facts, names and figures the model could not have guessed
Stock-phrase densityRate of cliché and hedging constructions per 1,000 wordsHigh. “It is important to note”, “plays a key role”Deletes the phrase and states the claim, or cuts the sentence

We have a baseline for one of those signals from our own measurement. Across 531 articles ranking for 40 B2B content queries (US, September 2026), median burstiness is 0.79. That is what ranking editorial pages produce, not a target invented by a tool vendor.

So how do AI detectors work? They score signals like these, fuse them into a single probability, and hand you a percentage. The math is sound. Trouble starts when the percentage is read as evidence.

Watermarks: the signal that is put there on purpose

A detector reads a footprint the model left by accident. A watermark is the opposite, a pattern the provider adds on purpose while the model samples tokens, invisible to the reader. Scott Aaronson proposed one in November 2022 while on leave at OpenAI, “a tool for statistically watermarking the outputs of a text model like GPT.” Kirchenbauer, Goldstein and four co-authors published the green-list scheme in January 2023: it selects “a randomized set of ‘green’ tokens” before each word, then softly promotes them during sampling.

Google DeepMind shipped it. Sumanth Dathathri and colleagues described SynthID-Text in Nature in October 2024, a method that works by “modifying the next-token sampling procedure.” Run on Gemini traffic against an unwatermarked model, it moved the thumbs-up rate by 0.01%, which the paper calls statistically insignificant. The code shipped in Hugging Face Transformers v4.46.0 that month, and the SynthID Detector portal opened to early testers in May 2025.

A watermark tells a platform where text came from. It tells an editor nothing about whether the text is any good.

Two providers then went opposite ways. Anthropic announced on August 14, 2026 that “Future Claude models will generate text that contains a watermark,” derived from SynthID-Text and applied “globally at launch.” Its help page now lists the watermark on current Claude models and says older ones will be “covered by December 2, 2026.” OpenAI has a text method it has not released, saying in an August 2024 update that it “is less robust against globalized tampering” such as translation, and that “it could stigmatize use of AI as a useful writing tool for non-native English speakers.” Its images and audio carry provenance marks today; for text, its help page states only a goal “to expand provenance signals to all modalities including text.”

ProviderWhat existsWho can checkStatus
Google DeepMindSynthID-Text, tested on live Gemini trafficGoogle, plus anyone using the released codeOctober 2024; tester portal May 2025
AnthropicA SynthID-Text watermark on current Claude modelsEligible organizations, in private previewLive; older models by December 2, 2026
OpenAIA text method built, not released; images and audio are markedNobody outside the company, for textText named a goal, July 2026

The timing is regulatory. Article 50(2) of the EU AI Act requires outputs “marked in a machine-readable format and detectable as artificially generated or manipulated” from August 2, 2026. A July 2026 amendment gives systems already on the market before that date until December 2, 2026. Anthropic, Google and OpenAI all signed the matching Code of Practice in July 2026.

For an editor the limits matter more than the launch. A watermark reports provenance, not craft. Anthropic says its key “can’t tell whether the text was written by a different AI,” and no cross-vendor check exists, so a public detector still runs on statistics. The mark also thins where writing gets routine: detection “doesn’t work well on small samples,” weakens under paraphrase, and is “sparser on factual passages.” The four signals above stay what an editor can act on. The watermark is a provenance fact a platform may use.

Why one detector is not proof

The best-documented failure is bias against non-native English writers. A Stanford study by Weixin Liang and colleagues ran seven off-the-shelf detectors over 91 human-written essays from TOEFL, the English test non-native applicants take for US universities. Every flag there is a false positive, human text called AI. The detectors “misclassified over half of the TOEFL essays as ‘AI-generated’ (average false positive rate: 61.22%),” and all seven unanimously flagged 18 of the 91. On essays by US eighth-graders, the same detectors were close to perfect.

The cause is mechanical. According to Tara García Mathewson, reporting for The Markup in August 2023, “AI detectors tend to be programmed to flag writing as AI-generated when the word choice is predictable and the sentences are more simple.” Non-native writers produce that profile. So does a good technical manual, a compliance page, and anyone writing in a second language for a client.

Plain, predictable English is what a detector flags, and plenty of careful human writers produce exactly that.

OpenAI ran into the ceiling on its own product. The company’s classifier launch page now opens with a note: “As of July 20, 2023, the AI classifier is no longer available due to its low rate of accuracy.” At retirement it “correctly identifies 26% of AI-written text (true positives) as ‘likely AI-written,’ while incorrectly labeling human-written text as AI-written 9% of the time.” Read the retirement notice before you quote any vendor’s accuracy claim.

OpenAI's own classifier at retirement
AI-written text it caught 26%
Human-written text it called AI 9%
Retired on July 20, 2023 for its low rate of accuracy. Source: OpenAI.

Detectors did get better after 2023. GPTZero retrained on writing by people using English as a second language and reports it “reduced the false positive rate on TOEFL essays to just 1.1%.” A rival, in its Pangram technical report, scores that same model against the same 91 essays and reports “a false positive rate of 7.7%, or 1.1% if ‘Possible AI Content’ is generously labeled a negative.” Both numbers are honest. They differ on one point, whether “possible AI” counts as an accusation. To the writer being accused, it does.

One wrong flag costs a freelancer a client. Price that before a score decides anything.

That gap is the argument for reading a score as a range, and for pricing what a wrong flag costs before you rely on one. The Markup’s reporting describes Turnitin labeling more than 90 percent of one Johns Hopkins student’s paper as AI-generated. Run a process that treats such a number as a verdict, and you drop a freelancer over a sentence-length statistic. So the working rule is narrow. Read every score as a range, and require a second signal before that range means anything. Never act on one score alone.

How to tell if text is AI written without running a tool

Read the draft for the same signals the classifier reads for, then fix the writing rather than the score. Swapping words does nothing, and the reason shows up clearly when the three states sit side by side:

VersionSentence rhythmParagraph openersWhat a reader notices
Raw AI outputUniform lengthsRepeats “The”, “This”, “It”Stock phrasing, claims with no source
Word-swapped outputStill uniformSame openers, new synonymsReads stiffer, same shape underneath
Structurally editedMixed short and longNames, numbers, questionsSpecific, quotable, easier to trust

The highest-value edit is not rhythm. It’s evidence, and the ranking field leaves the door wide open here. In those 531 ranking articles we measured a median of 0.9 statistics per 1,000 words, and 27% of them name zero sources at all. Among pages that do cite something, 22% cite a single host, and 13% carry at least one link returning 404 or 410 (only those two codes count as dead here; timeouts and blocked responses are reported separately).

One paragraph with a named source, a date and a number puts a draft ahead of most of what ranks. It also moves token predictability in the right direction, because a specific figure is exactly what a model would not have guessed. The limitation of our figures is worth stating: they describe pages that rank in the top 20 for those 40 queries, not the web at large.

A named source, a date and a number do more for a draft than any synonym swap.

Vocabulary matters less than the word lists suggest. Of those same ranking pages, 28% sit above our AI-vocabulary line, more than six high-severity tells per 1,000 words. More than a quarter of the ranking field trips the marker that most “beat the detector” advice is built on, which tells you the marker alone separates very little. Structure carries the rest, and the writing patterns that give away machine text run past vocabulary into paragraph and clause shape.

One more thing to settle before you wire a detector into a publishing workflow: detectors aren’t ranking systems. Ahrefs correlated AI-content share with position across 600,000 pages and got 0.011, “effectively zero,” and found “86.5% of top-ranking pages contain some amount of AI-generated content.” Its larger 2026 follow-up reads the other way, reporting that “higher AI use is correlated with lower ranking positions,” with pages under 50% AI content taking 82.2% of top-3 spots.

Neither study describes a detector gate. Our own corpus points the same way. Median statistics, burstiness and AI-tell rates barely move between positions 1 to 3 and positions 11 to 20. Google isn’t running your draft through a classifier. Readers still notice a flat one.

What the score is actually for

A detector score is a lead, not a verdict. It tells you where a draft reads mechanically, in the same way a readability grade tells you where a sentence runs long. What it can’t tell you is who wrote the text, and the research above shows what happens to people when it is asked to.

The work that improves the score is the work that improves the piece: vary the rhythm, kill repeated openers, replace one generic claim per section with a number and its source, and cut the phrases that survive only because nobody reread them.

Frequently asked

How do AI detectors work?

Statistical patterns, not words. The DetectGPT paper from Stanford identified the underlying property: model-generated text “tends to occupy negative curvature regions of the model’s log probability function.” Commercial tools turn that into a trained classifier. GPTZero, for example, runs “a sentence-by-sentence classification model” and reports a probability per sentence and per document.

Are AI detectors accurate?

Accuracy depends on whose benchmark you read. OpenAI retired its own classifier on July 20, 2023, “due to its low rate of accuracy,” after it caught 26% of AI text while flagging 9% of human text. Newer detectors do better: the Pangram technical report puts a retrained GPTZero model at 7.7% false positives on the TOEFL benchmark, or 1.1% when “Possible AI Content” is not counted as a flag.

Do AI detectors flag non-native English writers?

They did, heavily. Liang and colleagues found an average false positive rate of 61.22% across seven detectors on 91 TOEFL essays, with 18 of those essays flagged unanimously. The Markup traced the mechanism: detectors flag predictable word choice and simpler sentences, which describes a competent second-language writer as well as a model.

Does AI content hurt Google rankings?

No detector-style penalty shows up in the data. Ahrefs measured a 0.011 correlation between AI-content share and position across 600,000 pages, and found 86.5% of top-ranking pages contain some AI content. Its 2026 study of a larger sample found higher AI use correlated with lower positions, while pages under 50% AI content held 82.2% of top-3 spots. Quality separates them, not authorship.

Will paraphrasing an AI draft change the result?

Running a thesaurus over a draft changes vocabulary and leaves sentence-length variation and opener patterns untouched, so the structural signals stay where they were. The edit that moves them is the same edit that makes the piece worth reading: real evidence, varied rhythm, fewer stock phrases.

Can a watermark tell if text is AI written?

Only for one provider’s model, and only sometimes. Anthropic says its key answers how likely it is that Claude was involved, and that it “can’t tell whether the text was written by a different AI.” The Nature paper behind the method reports that longer texts carry more watermarking evidence and that paraphrasing weakens the mark. OpenAI has not released one for text. A watermark is a provenance signal, not a verdict on the writing, and nothing reads across vendors.

Sources

22 sources

Methodology: item 1 is our own measurement, and its rules are stated with it. Items 2 to 22 are primary sources: four arXiv papers, one peer-reviewed journal article, ten provider and vendor pages, one news investigation, one university magazine, one researcher’s blog, two EU legal texts and one European Commission announcement. The sources were first checked against the page they are quoted from on September 8 and September 16, 2026; every one was fetched again, and five were added, on September 29, 2026.

  1. 01Avoid Content corpus measurement, September 2026. The instrument we ran over the corpus covered 531 articles ranking in Google organic top-20 for 40 B2B content-marketing queries (US, English), SERP pulled September 7, 2026. Dead links counted as HTTP 404 and 410 only.Supports: every "we measured" figure above. Dataset not yet public, thresholds stated inline.
  2. 02Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., Finn, C. “DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature.” arXiv:2301.11305. https://arxiv.org/abs/2301.11305.Supports: the probability-curvature property detectors are built on.
  3. 03Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., Zou, J. “GPT detectors are biased against non-native English writers.” arXiv:2304.02819. https://arxiv.org/abs/2304.02819.Supports: 61.22% average false positive rate over 91 TOEFL essays, 18 of them flagged unanimously by all seven detectors.
Show all 22 sources
  1. 04OpenAI. “New AI classifier for indicating AI-written text.” https://openai.com/index/new-ai-classifier-for-indicating-ai-written-text/.Supports: retirement on July 20, 2023 "due to its low rate of accuracy", 26% true positives and 9% false positives.
  2. 05GPTZero. “Technology.” https://gptzero.me/technology.Supports: sentence-by-sentence classification, 1.1% false positive rate on TOEFL essays after retraining.
  3. 06Princeton Alumni Weekly. “Wrap Your Brain Around This.” February 24, 2023. https://paw.princeton.edu/article/wrap-your-brain-around.Supports: GPTZero as Edward Tian's senior thesis project and perplexity and burstiness as the two quantities he named.
  4. 07Emi, B., Spero, M. “Technical Report on the Pangram AI-Generated Text Classifier.” arXiv:2402.14873. https://arxiv.org/abs/2402.14873.Supports: 7.7% (or 1.1%) false positive rate for the retrained GPTZero model on the same TOEFL benchmark.
  5. 08García Mathewson, T. “AI Detection Tools Falsely Accuse International Students of Cheating.” The Markup, August 14, 2023. https://themarkup.org/machine-learning/2023/08/14/ai-detection-tools-falsely-accuse-international-students-of-cheating.Supports: why detectors flag predictable, simpler prose, and the Turnitin 90% case.
  6. 09Ong, S. Q. “AI-Generated Content Does Not Hurt Your Google Rankings (600,000 Pages Analyzed).” Ahrefs, July 7, 2025. https://ahrefs.com/blog/ai-generated-content-does-not-hurt-your-google-rankings/.Supports: the 0.011 correlation, and 86.5% of top-ranking pages containing some AI content.
  7. 10Law, R. “Google Doesn’t Punish AI Content; It Punishes Bad Content.” Ahrefs, July 27, 2026. https://ahrefs.com/blog/google-doesnt-punish-ai-content/.Supports: 1,000,000 pages sampled from the top 10 of 100,000 SERPs (June 2026), about 300,000 available to the crawler and 150,000 with enough content for a detection pass, with pages under 50% AI content holding 82.2% of top-3 rankings.
  8. 11Aaronson, S. “My AI Safety Lecture for UT Effective Altruism.” Shtetl-Optimized, November 2022. https://scottaaronson.blog/?p=6823.Supports: the 2022 watermarking proposal made while the author was on leave at OpenAI, quoted verbatim.
  9. 12Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., Goldstein, T. “A Watermark for Large Language Models.” arXiv:2301.10226, ICML 2023. https://arxiv.org/abs/2301.10226.Supports: the green-list scheme that promotes a randomized set of tokens during sampling.
  10. 13Dathathri, S., et al. “Scalable watermarking for identifying large language model outputs.” Nature, October 23, 2024. https://www.nature.com/articles/s41586-024-08025-4.Supports: SynthID-Text modifying the next-token sampling procedure, the 0.01% thumbs-up difference in the Gemini live experiment, and the stated limits on short texts, paraphrasing and low-entropy output.
  11. 14Google DeepMind and Hugging Face. “Introducing SynthID Text.” Hugging Face blog, October 23, 2024. https://huggingface.co/blog/synthid-text.Supports: the release of SynthID Text in Transformers v4.46.0.
  12. 15Google. “SynthID Detector, a new portal to help identify AI-generated content.” The Keyword, May 20, 2025. https://blog.google/innovation-and-ai/products/google-synthid-ai-content-detector/.Supports: the verification portal for content made with Google AI, rolled out to early testers first.
  13. 16Anthropic. “How Claude’s text watermarking works.” August 14, 2026. https://www.anthropic.com/news/claude-text-watermark.Supports: future Claude models carrying a SynthID-Text-derived watermark, the global rollout, the July 2026 EU Code of Practice signature, and the limits on small samples, other vendors' text and factual passages.
  14. 17Anthropic. “How Claude marks AI-generated content.” Claude Help Center. https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content.Supports: the watermark on current Claude models, the December 2, 2026 date for older models, and detection limited to eligible organizations.
  15. 18OpenAI. “Understanding the source of what we see and hear online.” May 2024, updated August 4, 2024. https://openai.com/index/understanding-the-source-of-what-we-see-and-hear-online/.Supports: a text watermarking method built and not released, its weakness against translation and rewording, and the stated risk to non-native English speakers.
  16. 19OpenAI. “Provenance signals (Content Credentials, SynthID) in OpenAI-generated content.” OpenAI Help Center. https://help.openai.com/en/articles/8912793-provenance-signals-content-credentials-synthid-in-openai-generated-content.Supports: provenance marks on images and audio, none on text, and the stated goal of extending them to text.
  17. 20European Union. “Article 50: Transparency Obligations for Providers and Deployers of Certain AI Systems.” EU AI Act. https://artificialintelligenceact.eu/article/50/.Supports: the machine-readable marking requirement of Article 50(2) and its application from August 2, 2026.
  18. 21European Union. Regulation (EU) 2026/1744 amending the AI Act, done at Strasbourg on July 8, 2026. https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:32026R1744.Supports: the new Article 111(4) giving systems placed on the market before August 2, 2026 until December 2, 2026 to comply with Article 50(2).
  19. 22European Commission. “Strong backing for the Code of Practice on Transparency of AI-generated Content.” Shaping Europe’s digital future. https://digital-strategy.ec.europa.eu/en/news/strong-backing-code-practice-transparency-ai-generated-content.Supports: Anthropic, Google and OpenAI among the signatories, and about 190 signatories by the end of July 2026.

Research · 8 min read

Every article, with its receipt.

One article free on sign-up. No credit card. Agencies start with a pilot.