Skip to main content
v2026.11,772 entries · CC-BY 4.0
Dictionary termTrack AStablev2026.2

Detection tool (AI-generated)

A software system that estimates the probability that a given piece of content -- typically text, sometimes images -- was produced by a generative AI system, usually by statistically analysing patterns in the content itself without any cooperation from or signal embedded by the original generator, as distinct from watermarking, which requires generator-side cooperation at the point of creation.

ByCASRAI Editorial Board
· Last updated 22 Aug 2026
Share this

Ask CASRAI · included with Regulatory Radar

Ask about Detection tool (AI-generated)

Ask CASRAI answers research-administration questions and cites the passages behind every claim — and says so when the corpus does not cover something, instead of guessing. It comes with a Regulatory Radar subscription at $29 a month, alongside the daily digest of regulatory changes and the dashboard of what changed.

150 questions a day, on this site, over the API, or inside your own tools through the CASRAI MCP server.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Examples

Worked examples

  • Is an instance

    An institutionally-licensed detector (e.g. Turnitin's AI writing detection module) flagging a submitted student paper as likely containing AI-generated text, routed to a human for review rather than an automatic penalty

  • Is an instance

    A publisher-side integrity screening tool (e.g. Springer Nature's Geppetto, now part of the STM Integrity Hub) flagging a manuscript section for further human investigation

Counter-examples

Looks similar, but isn't

  • Not an instance

    A tool that verifies an embedded, generator-created watermark is watermarking verification, not the forensic pattern-based inference this entry covers

  • Not an instance

    A human reviewer's subjective impression that a document 'reads like AI' with no software tool involved is not a detection tool in this sense

Editorial commentary

These tools infer likely AI origin from statistical patterns — unusually uniform sentence structure, low ‘perplexity’ or predictability, characteristic phrasing — rather than from any cooperative signal deliberately embedded by a generator. That inferential, probabilistic basis is the source of their central, well-documented limitation, and it needs to be stated plainly here because overclaiming detector reliability causes real harm to accused students and researchers.

The false-positive problem is real, documented, and disproportionate

Vendor-claimed accuracy figures are consistently higher than what independent testing finds. Turnitin has publicly claimed roughly 98% accuracy with a false-positive rate under 1% for documents where more than 20% of the text is AI-generated; independent 2024-2025 testing has generally found lower accuracy on unedited generative-AI output and false-positive rates climbing meaningfully higher — reported in some studies in the range of 5-12% — specifically on non-native-English writing, heavily edited drafts, and technical prose. OpenAI’s own AI Text Classifier, launched January 2023, was withdrawn by OpenAI itself in July 2023, citing a ‘low rate of accuracy’: OpenAI’s own testing found it correctly identified only about 26% of AI-written text while mislabelling roughly 9% of genuinely human-written text as AI-generated.

The consequences are not hypothetical. Orion Newby, an Adelphi University student with learning/neurological disabilities, was accused of academic dishonesty after an AI-detection flag; a New York state court ultimately ruled in his favour, reversed the disciplinary findings, and ordered his record expunged. Vanderbilt University disabled Turnitin’s AI-writing detection feature in 2023 over false-positive risk at scale; Curtin University announced it would disable the same feature from January 2026, citing reliability concerns. Several institutions have stopped using these tools, or use them only as a prompt for human conversation rather than as evidence of misconduct on their own. See why does my paper say AI-detected? and Turnitin AI detection vs. standalone AI detectors for more detail, and treat any single detector flag as a starting point for human review, never as standalone proof.

How this differs from related AI-band terms

  • vs. watermarking: detection is forensic and after-the-fact, working (unreliably) on any content regardless of source cooperation; watermarking is proactive and requires the generator to have embedded a signal, which makes it more reliable when present but inapplicable when the generator didn’t cooperate.
  • vs. AI provenance: provenance is the broader documentation goal; detection tools are one (unreliable) attempt to reconstruct provenance information that was never actually recorded.

Also known as

AI detector · GPT detector

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="Detection tool (AI-generated)"
      vocab-term-identifier="https://casrai.org/dictionary/term/detection-tool-ai-generated" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/detection-tool-ai-generated",
  "name": "Detection tool (AI-generated)",
  "identifier": "https://casrai.org/dictionary/term/detection-tool-ai-generated",
  "description": "A software system that estimates the probability that a given piece of content -- typically text, sometimes images -- was produced by a generative AI system, usually by statistically analysing patterns in the content itself without any cooperation from or signal embedded by the original generator, as distinct from watermarking, which requires generator-side cooperation at the point of creation.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/genai-disclosure#set",
  "url": "https://casrai.org/dictionary/term/detection-tool-ai-generated",
  "sameAs": [
    "AI detector",
    "GPT detector"
  ],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "author": {
    "@id": "https://casrai.org/#editorial-team"
  },
  "datePublished": "2026-05-21T01:55:32",
  "dateModified": "2026-08-22T14:52:05",
  "inLanguage": "en-GB",
  "isAccessibleForFree": true
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 72,264 indexed passages, and every answer cites the ones it drew on.