The people behind AI text generation, detection and watermarking

Detection and watermarking claims usually cite a lab or a university. Here are the people behind them, one result each, with the limits they published.

Written by Avoid Content TeamSep 29, 20268 min read

Read the name before you read the claim

A tool page that says “based on Stanford research” is not citing one thing. Two Stanford papers from 2023 sit behind most of those claims and they answer different questions. One describes a statistical property of model output. The other reports what seven detectors did to essays by non-native English speakers. A reader who stops at the university name cannot tell which.

Three checks turn a name into information. Find the paper the claim points to. Read its scope line, the sentence that says what was measured and on what. Check the date, because a result about 2023 detectors is a statement about 2023.

A lab's name on a tool page points to a paper, and the claim lives in that paper's scope line.

The people behind AI text generation, detection and watermarking are a short list, and their work is public. Each entry below gives the name, the affiliation at the time, the year, one result and what it changes for an editor.

Generation

Ashish Vaswani and seven co-authors: the Transformer (2017)

Eight researchers posted “Attention Is All You Need” to arXiv on June 12, 2017. Six of them listed a Google Brain or Google Research address on the paper, one the University of Toronto. Their claim is in the abstract: “We propose a new simple network architecture, the Transformer, based solely on attention mechanisms, dispensing with recurrence and convolutions entirely.”

For an editor, this is the design under every model that now writes marketing copy, and the reason detection research talks about token probabilities instead of suspicious words.

One 2017 architecture and one training goal, guessing the next word, sit under all of today's copywriting models.

Alec Radford, Ilya Sutskever and co-authors: GPT and GPT-2 (2019)

OpenAI turned that architecture into a recipe in two reports on its own site. The first, by Alec Radford, Karthik Narasimhan, Tim Salimans and Ilya Sutskever, set out generative pre-training on unlabeled text, then fine-tuning per task. The second, by Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei and Sutskever, all listed at OpenAI, reports that “language models begin to learn these tasks without any explicit supervision when trained on a new dataset of millions of webpages called WebText.” OpenAI’s announcement calls GPT-2 “a successor to GPT.”

An editor should read the objective: the model predicts the next word over web text, so its default draft is the middle of the web on your topic.

Detection

Detection is a separate line of work. The mechanics sit in our guide to how AI detectors work; what follows is who produced the results it rests on.

Sebastian Gehrmann, Hendrik Strobelt and Alexander Rush: GLTR (2019)

Sebastian Gehrmann and Alexander Rush of Harvard’s engineering school and Hendrik Strobelt of IBM Research and the MIT-IBM Watson AI lab presented GLTR at ACL 2019 as “a tool to support humans in detecting whether a text was generated by a model.” It “applies a suite of baseline statistical methods that can detect generation artifacts across multiple sampling schemes,” and in the authors’ human-subjects study it “improves the human detection-rate of fake text from 54% to 72% without any prior training.”

People spotting generated text, with and without GLTR
Without the tool 54%
With GLTR's highlighting 72%
Human-subjects study reported by Gehrmann, Strobelt and Rush at ACL 2019.

For an editor, that is the useful framing: a statistical reading helps a person look, and the person still decides.

OpenAI: a detector of its own, twice (2019 and 2023)

OpenAI shipped a detector with the full GPT-2 model on November 5, 2019: a classifier with “detection rates of ~95% for detecting 1.5B GPT-2-generated text.” The same post added a warning: “We believe this is not high enough accuracy for standalone detection and needs to be paired with metadata-based approaches, human judgment, and public education to be more effective.” A second attempt, the AI classifier of January 31, 2023, was withdrawn on July 20, 2023 “due to its low rate of accuracy,” after catching 26% of AI-written text and labeling human text as AI 9% of the time.

The company that built the models said twice, in writing, that a detector score should not stand on its own. An editor can quote that before anyone else does.

Eric Mitchell and Chelsea Finn: DetectGPT (2023)

Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning and Chelsea Finn, all at Stanford University, posted DetectGPT to arXiv on January 26, 2023, published at ICML 2023.

The paper is open access: DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature (Mitchell, Lee, Khazatsky, Manning and Finn, 2023, arXiv:2301.11305, CC BY 4.0). Its finding: “text sampled from an LLM tends to occupy negative curvature regions of the model’s log probability function.” The method reads log probabilities and perturbed copies of a passage, with no trained classifier behind it.

That is why a thesaurus pass moves nothing for an editor: the signal is a property of how models pick words, not a list of words.

Edward Tian: GPTZero (2023)

Princeton Alumni Weekly reported on February 24, 2023 that Edward Tian, then a computer science undergraduate in the class of 2023, had developed “a program called GPTZero, which purports to detect whether a piece of writing was generated by artificial intelligence” as his senior thesis project. The article names the two quantities the program scores, which Tian called perplexity and burstiness: how predictable the words are and how much sentence length varies.

For an editor, those two words are the vocabulary nearly every detector report still uses to explain a score.

A student project from early 2023 still supplies the words most AI scores are explained in.

Weixin Liang and James Zou: the detector bias result (2023)

Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou, all at Stanford University, ran seven widely used detectors over 91 human-written TOEFL essays, the English exam international applicants take for US universities. Every flag there is a false positive. The paper reports that “they misclassified over half of the TOEFL essays as ‘AI-generated’ (average false positive rate: 61.22%),” while essays by US eighth-graders were scored close to perfectly.

Quote that scope line before a score is treated as evidence about a person: seven products, that corpus, 2023.

Tara García Mathewson: what a false flag costs (2023)

Reporting for The Markup on August 14, 2023, Tara García Mathewson followed the Stanford finding into classrooms and gave the mechanism in one line: “AI detectors tend to be programmed to flag writing as AI-generated when the word choice is predictable and the sentences are more simple.”

This predicts who on an editor’s team gets flagged: second-language writers, and anyone writing plain technical copy.

A detector flag describes how predictable the prose is. Who wrote it is a separate question the score cannot answer.

Watermarking

A detector reads a footprint the model left by accident. A watermark is a pattern the provider adds on purpose while the model samples words. What that changes for a published draft sits in the watermark section of our detector guide.

Scott Aaronson: the proposal (2022)

In a lecture posted to his blog in November 2022, Scott Aaronson described what he was working on while on leave at OpenAI: “My main project so far has been a tool for statistically watermarking the outputs of a text model like GPT.” The goal he states is an unnoticeable signal in the model’s word choices that can be checked later.

For an editor, this is the origin of the idea that provenance can be built into text rather than inferred afterwards.

John Kirchenbauer and Tom Goldstein: the green list (2023)

John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers and Tom Goldstein, all at the University of Maryland, posted “A Watermark for Large Language Models” on January 24, 2023 and published it at ICML 2023. The scheme works by “selecting a randomized set of ‘green’ tokens before a word is generated, and then softly promoting use of green tokens during sampling.” A statistical test then reports whether a passage carries the mark.

Editors get the first published method precise enough to argue about: detectable without model access, weakened by rewriting.

Rewriting a watermarked passage thins the mark, so a clean result on edited text proves little.

Sumanth Dathathri and Pushmeet Kohli: SynthID-Text (2024)

Nature published “Scalable watermarking for identifying large language model outputs” on October 23, 2024, with authors at Google DeepMind. The method “works by carefully modifying the next-token sampling procedure to inject subtle, context-specific modifications into the generated text distribution.” Google DeepMind’s own post says the text project “was led by Sumanth Dathathri and Pushmeet Kohli.” The paper reports a live test on Gemini traffic where the thumbs-up rate between the watermarked and unwatermarked model differed by 0.01%, which the authors call statistically insignificant.

Watermarking stopped being a proposal here and started running in a product an editor may already use.

Anthropic and OpenAI: two decisions about shipping a watermark (2024 and 2026)

Two providers went opposite ways. Anthropic wrote on August 14, 2026 that “Future Claude models will generate text that contains a watermark,” describing it as a version of the SynthID-Text approach Google DeepMind published in Nature, applied to comply with the EU AI Act. By the end of September its help page listed the watermark on its current models, with older models to be “covered by December 2, 2026.” OpenAI wrote in an August 2024 update that “Our teams have developed a text watermarking method that we continue to consider as we research alternatives,” citing weakness against translation and the risk of stigmatizing AI use by non-native English speakers. Its images and audio now carry provenance marks; for text, its help page names a goal “to expand provenance signals to all modalities including text” under the EU code it signed in July 2026, and nothing has shipped.

The consequence for an editor is narrow: a watermark answers a question about one company’s model, and nothing reads across vendors.

A watermark names one vendor's model. Whether the draft deserves to run is still an editor's call.

Our reading

Three things follow, and they are our reading rather than any paper’s finding.

Detection measures text, not authorship. Every result above describes the statistics of a passage, and the strongest published evaluation of the class is a false-positive rate on human writing. We treat a score as a lead, never as proof, and we set the two questions side by side in our piece on detectors against a citation check.

Watermarking is provenance, and it belongs to the provider. It answers whether one company’s model was involved. Anthropic’s version arrived because a regulation required it, and no key reads another vendor’s text. That is useful to a platform and says nothing about whether a draft is worth publishing.

Neither one grades the writing. Evidence does, and evidence is scarce: across 531 articles ranking for 40 B2B content queries (US, September 2026), 27% name zero sources. An editor closes that gap by naming a source for each claim and confirming it survives a real verification pass. The researchers on this page published the limits of their own methods, and reading those limits beats arguing with a percentage.

The people and the results

WhoYearResultFor an editor
Vaswani and seven co-authors, Google and Toronto2017The Transformer, “based solely on attention mechanisms”The design under every model that drafts copy
Radford, Sutskever and co-authors, OpenAI2019GPT-2 learns tasks “without any explicit supervision” from web textThe default draft is the middle of the web
Gehrmann, Strobelt and Rush, Harvard and IBM2019GLTR raised human detection of generated text from 54% to 72%A statistical aid helps a person look
OpenAI2019, 2023A GPT-2 detector it judged too weak to stand alone; a 2023 classifier withdrawnThe model maker’s own warning about scores
Mitchell, Finn and co-authors, Stanford2023DetectGPT: model text sits in “negative curvature regions”Swapping words changes nothing
Edward Tian, Princeton2023GPTZero scores perplexity and burstinessThe two terms detector reports still use
Liang, Zou and co-authors, Stanford2023Seven detectors, 91 TOEFL essays, 61.22% average false positivesQuote the scope before a score judges a person
Tara García Mathewson, The Markup2023Detectors flag “predictable” and “more simple” writingSecond-language and plain technical writers get flagged
Scott Aaronson, on leave at OpenAI2022A proposal to watermark model output statisticallyProvenance can be built in
Kirchenbauer, Goldstein and co-authors, Maryland2023The green-list watermark, tested with p-valuesDetectable without the model, weakened by rewriting
Dathathri, Kohli and team, Google DeepMind2024SynthID-Text on live Gemini traffic, 0.01% thumbs-up differenceWatermarking became a shipped feature
OpenAI2024, 2026A text watermark built and not released; text named a goal in 2026Nobody outside can check its text
Anthropic2026A watermark on current Claude models, older ones by December 2, 2026One vendor’s provenance, no cross-vendor check

Frequently asked

Who invented AI detectors?

No single person, on the record we checked. Public tools go back to 2019: GLTR, from Sebastian Gehrmann and Alexander Rush at Harvard and Hendrik Strobelt at IBM Research, and the detector OpenAI released with GPT-2. Two 2023 results carry most of today’s citations: DetectGPT, from five authors at Stanford University, which detects model text from probability curvature without a classifier, and GPTZero, reported by Princeton Alumni Weekly in February 2023 as Edward Tian’s senior thesis project.

What did the DetectGPT paper actually show?

That machine-generated text sits in a measurable place. The paper states that “text sampled from an LLM tends to occupy negative curvature regions of the model’s log probability function,” and turns that into a detection criterion using only log probabilities and perturbed copies of the passage. No classifier is trained on collected AI text.

Why do detectors flag writing by non-native English speakers?

Because the profile they score is the profile of careful second-language writing. Liang and four Stanford co-authors measured an average false positive rate of 61.22% across seven detectors on 91 human-written TOEFL essays. The Markup gave the mechanism: detectors “flag writing as AI-generated when the word choice is predictable and the sentences are more simple.”

Does a watermark prove a text was written by AI?

It reports provenance for one provider. Anthropic states that its key answers how likely it is that Claude was involved and “can’t tell whether the text was written by a different AI.” The Nature paper behind the method notes that longer passages carry more evidence and that paraphrasing weakens the mark. There is no cross-vendor check.

Sources

20 sources

Methodology: item 1 is our own measurement, and its rule is stated with it. Items 2 to 20 are primary sources: four arXiv papers, one conference paper, one peer-reviewed journal article, ten company pages, reports and help pages, one news investigation, one researcher’s blog and one university magazine. Each was fetched and the quoted sentence located on the page on September 16, 2026; all of them were fetched again, and items 6 to 8, 18 and 20 added, on September 29, 2026.

  1. 01Avoid Content corpus measurement, September 2026. The instrument we ran over the corpus covered 531 articles ranking in Google organic top-20 for 40 B2B content-marketing queries (US, English), SERP pulled September 7, 2026.Supports: 27% of those pages naming zero sources. Dataset not yet public; the figure describes pages ranking for those 40 queries, not the web at large.
  2. 02Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., Polosukhin, I. “Attention Is All You Need.” arXiv:1706.03762, submitted June 12, 2017. https://arxiv.org/abs/1706.03762.Supports: the Transformer proposal, the author list and the Google Brain, Google Research and University of Toronto affiliations on the paper's first page.
  3. 03Radford, A., Narasimhan, K., Salimans, T., Sutskever, I. “Improving Language Understanding by Generative Pre-Training.” OpenAI. https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf.Supports: generative pre-training followed by fine-tuning, and the four authors listed at OpenAI.
Show all 20 sources
  1. 04Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I. “Language Models are Unsupervised Multitask Learners.” OpenAI. https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf.Supports: the WebText sentence quoted above and the OpenAI affiliation of all six authors.
  2. 05OpenAI. “Better language models and their implications.” https://openai.com/index/better-language-models/.Supports: GPT-2 as "a successor to GPT", the original-post author list including Alec Radford and Ilya Sutskever, and the May 2019 interim update that dates the release year.
  3. 06Gehrmann, S., Strobelt, H., Rush, A. M. “GLTR: Statistical Detection and Visualization of Generated Text.” ACL 2019, system demonstrations. https://aclanthology.org/P19-3019/.Supports: the Harvard and IBM Research affiliations, the tool's stated purpose, the quoted method sentence and the 54% to 72% human detection-rate result.
  4. 07OpenAI. “GPT-2: 1.5B release.” November 5, 2019. https://openai.com/index/gpt-2-1-5b-release/.Supports: the released detector, its ~95% detection rate on 1.5B GPT-2 text, and the quoted warning against standalone detection.
  5. 08OpenAI. “New AI classifier for indicating AI-written text.” January 31, 2023. https://openai.com/index/new-ai-classifier-for-indicating-ai-written-text/.Supports: the withdrawal on July 20, 2023 "due to its low rate of accuracy", 26% true positives and 9% false positives.
  6. 09Mitchell, E., Lee, Y., Khazatsky, A., Manning, C. D., Finn, C. “DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability Curvature.” arXiv:2301.11305, ICML 2023. https://arxiv.org/abs/2301.11305.Supports: the negative-curvature property, the no-classifier method, the January 26, 2023 submission date, the Stanford University affiliation on the paper's first page and its CC BY 4.0 license.
  7. 10Princeton Alumni Weekly. “Wrap Your Brain Around This.” February 24, 2023. https://paw.princeton.edu/article/wrap-your-brain-around.Supports: Edward Tian, class of 2023, computer science undergraduate, GPTZero as a senior thesis project, and the perplexity and burstiness pair.
  8. 11Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., Zou, J. “GPT detectors are biased against non-native English writers.” arXiv:2304.02819. https://arxiv.org/abs/2304.02819.Supports: seven detectors, 91 human-authored TOEFL essays, the 61.22% average false positive rate and the Stanford University affiliations.
  9. 12García Mathewson, T. “AI Detection Tools Falsely Accuse International Students of Cheating.” The Markup, August 14, 2023. https://themarkup.org/machine-learning/2023/08/14/ai-detection-tools-falsely-accuse-international-students-of-cheating.Supports: the mechanism quote, the byline and the August 14, 2023 date.
  10. 13Aaronson, S. “My AI Safety Lecture for UT Effective Altruism.” Shtetl-Optimized, November 2022. https://scottaaronson.blog/?p=6823.Supports: the statistical watermarking proposal, quoted verbatim, and the author's own statement that he was on leave at OpenAI.
  11. 14Kirchenbauer, J., Geiping, J., Wen, Y., Katz, J., Miers, I., Goldstein, T. “A Watermark for Large Language Models.” arXiv:2301.10226, submitted January 24, 2023, ICML 2023. https://arxiv.org/abs/2301.10226.Supports: the green-list scheme, the p-value test, and the University of Maryland affiliation on the paper's first page.
  12. 15Dathathri, S., et al. “Scalable watermarking for identifying large language model outputs.” Nature, October 23, 2024. https://www.nature.com/articles/s41586-024-08025-4.Supports: the next-token sampling quote, the Google DeepMind affiliation, the October 23, 2024 publication date and the 0.01% thumbs-up difference in the live Gemini experiment.
  13. 16Google DeepMind. “Watermarking AI-generated text and video with SynthID.” https://deepmind.google/blog/watermarking-ai-generated-text-and-video-with-synthid/.Supports: the text watermarking project being led by Sumanth Dathathri and Pushmeet Kohli.
  14. 17Anthropic. “How Claude’s text watermarking works.” August 14, 2026. https://www.anthropic.com/news/claude-text-watermark.Supports: future Claude models carrying a watermark, the SynthID-Text lineage, the EU requirement, and the statement that the key cannot tell whether other AI wrote the text.
  15. 18Anthropic. “How Claude marks AI-generated content.” Claude Help Center. https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content.Supports: the Claude models carrying text watermarks, the December 2, 2026 date for older models, and detection limited to eligible organizations.
  16. 19OpenAI. “Understanding the source of what we see and hear online.” May 2024, updated August 4, 2024. https://openai.com/index/understanding-the-source-of-what-we-see-and-hear-online/.Supports: a text watermarking method built and not released (in the August 4, 2024 update), its weakness against translation, and the stated risk to non-native English speakers.
  17. 20OpenAI. “Provenance signals (Content Credentials, SynthID) in OpenAI-generated content.” OpenAI Help Center. https://help.openai.com/en/articles/8912793-provenance-signals-content-credentials-synthid-in-openai-generated-content.Supports: provenance marks on images and audio, none on text, and the stated goal of extending them to text under the EU Code of Practice.

Research · 8 min read

Every article, with its receipt.

One article free on sign-up. No credit card. Agencies start with a pilot.