Examples
Worked examples
- Is an instance
A manuscript flagged by PPS's tortured-phrases detector for consistently writing "haze figuring" instead of "cloud computing" and "profound learning" instead of "deep learning," then posted to PubPeer for author response before any journal investigation begins.
- Is an instance
A paper flagged under 'Feet of Clay' for citing an unusually high proportion of references that were later found to be retracted, prompting an editor to check whether the citations were substantive reliance or incidental.
Counter-examples
Looks similar, but isn't
- Not an instance
A paper using genuinely new, clearly-defined terminology proposed by an emerging subfield is not a tortured-phrase hit — the test is whether domain experts recognize it as a garbled substitution for an existing standard term, not a deliberate coinage.
- Not an instance
A paper citing a retracted study specifically to critique or discuss its retraction is not evidence of citation manipulation, even though it appears in a 'Feet of Clay'-style citation count.
Editorial commentary
The Problematic Paper Screener (PPS) is a public, automated screening website that scans the published literature for statistical and textual signatures associated with paper mills, plagiarism, and citation manipulation, then posts its findings for community and editorial review rather than issuing final determinations itself.
Who built and maintains it
The PPS was built and is maintained by Guillaume Cabanac, a computer scientist at the University of Toulouse (IRIT lab), in collaboration with Cyril Labbé and Alexander Magazinov. The site went live on 27 February 2021 and has been updated continuously since; it runs as an independent academic project, not a commercial product, and its detection logic and detector-by-detector output counts are published openly on the site itself.
What it detects
PPS combines several independent detectors, each aimed at a different signature of a manufactured or manipulated paper:
- Tortured phrases — garbled synonym substitutions for established technical terms (for example, “counterfeit consciousness” for artificial intelligence), the signature left by running text through paraphrasing tools to evade plagiarism checkers. See CASRAI’s tortured phrases entry for the full mechanism.
- Computer-generated papers — text produced by nonsense-generating programs such as SCIgen (computer science) and Mathgen (mathematics), which assemble grammatically plausible but meaningless technical prose.
- Feet of Clay — papers that cite an unusually large share of retracted, withdrawn, or expression-of-concern references, a pattern associated with citation manipulation and paper-mill output.
- Citation and reference-integrity issues more broadly, including papers citing works flagged elsewhere in the retraction ecosystem.
How it works
PPS harvests bibliographic and full-text metadata from Crossref (which incorporates the Retraction Watch Database), Dimensions, PubMed, and PubPeer, then applies pattern-matching and statistical checks to flag candidate papers. Rather than adjudicating misconduct, the tool generates screening reports that are posted publicly to PubPeer, where authors, editors, and other researchers can respond, explain, or dispute the flag before any institutional process begins.
Why the signal works
The underlying logic is statistical improbability. Genuine domain writing overwhelmingly uses established technical vocabulary; a paper that consistently substitutes an odd synonym for a well-known term is far more likely to be the output of an automated paraphrasing pass — used either to disguise plagiarized source text or to mass-produce manuscripts — than an author’s genuine word choice. Citation-based detectors work on a similar logic: heavy reliance on retracted sources, or coordinated citation to a narrow set of papers, is a pattern real scholarship rarely produces by accident.
Known false-positive modes
A PPS flag is evidence, not proof, and each detector has genuine failure modes that a reviewer has to rule out before treating a hit as meaningful:
- Tortured phrases: legitimate new terminology proposed in good faith by an emerging subfield, or unusual but consistent phrasing from a non-native English author working without a paraphrasing tool, can superficially resemble a tortured phrase. The distinguishing test is whether domain experts recognize the phrase as a garbled substitution for an already-standard term, not a deliberate, clearly-defined coinage.
- Feet of Clay: a paper can legitimately cite a retracted study to critique it, to describe research history, or because the citation predates the retraction and was never updated — none of that is citation manipulation.
- Computer-generated-text detectors: can occasionally mis-flag genuinely eccentric but human-authored technical prose, particularly in translated or heavily edited manuscripts.
For this reason PPS output functions as a triage signal for further human review, not a self-executing verdict.
How integrity officers and editors actually use it
A PPS or PubPeer-posted flag is the start of a screening workflow, not the end of one, and the distance between “screening hit” and “misconduct finding” is the whole point of a COPE-aligned process:
- Triage: an editor or research-integrity officer checks whether the flagged pattern actually holds up on manual reading — is the phrase genuinely a garbled standard term, is the retracted-citation share genuinely anomalous for the field and article type.
- Author contact: if the pattern survives triage, standard practice (consistent with COPE’s guidance on handling allegations) is to give the author an opportunity to explain before escalating — this is a screening inquiry, not yet a formal allegation of misconduct.
- Escalation only on unresolved concern: if the explanation doesn’t account for the pattern, or the pattern recurs across multiple submissions from related authors, the case moves into a formal investigation under the journal’s or institution’s misconduct policy, distinct from and following the initial screen. See CASRAI’s guide on research misconduct and ORI findings for how that later stage works.
Screening tools like PPS exist specifically to surface candidates for this process at a scale no editorial team could review manually; they are not a substitute for it.
Limitations
- Coverage is bounded by what Crossref, Dimensions, PubMed, and PubPeer index — papers outside those sources, or in venues with thin metadata, are not screened.
- Tortured-phrase detection depends on a maintained dictionary of known substitution pairs; paraphrasing methods (including newer generative-AI tools) continue to evolve, and detection can lag novel substitution patterns.
- PPS does not itself screen figures or images — image-integrity issues such as inappropriate duplication or manipulation are a separate detection problem, covered by tools discussed in CASRAI’s image duplication and image manipulation entries and the detecting image manipulation in figures guide.
- As an independent, largely volunteer-run project, PPS operates alongside — not as a replacement for — publisher-side infrastructure such as the STM Integrity Hub and coordinated initiatives like United2Act.
Frequently asked questions
Is a Problematic Paper Screener flag proof that a paper is fraudulent?
No. It is a screening signal that surfaces candidates for human review. Confirming misconduct requires editorial or institutional investigation, typically following a COPE-aligned process, not the automated flag alone.
Who can see PPS results?
PPS is a public website, and its screening reports are posted openly to PubPeer, where anyone — including the paper’s authors — can view and respond to them.
Does PPS replace publisher screening tools like the STM Integrity Hub?
No. PPS is an independent academic project focused on specific textual and citation signatures. Publisher-run infrastructure such as the STM Integrity Hub pools a broader set of pre-publication checks (duplicate submission, papermill risk scoring, retracted-reference cross-checking) across member publishers before peer review; the two are complementary, not competing.
Does the Problematic Paper Screener detect image manipulation?
No. Its detectors target text and citation patterns (tortured phrases, computer-generated prose, retracted-reference reliance). Image-integrity screening is handled by separate tools; see CASRAI’s image duplication entry for that detection landscape.
Machine-readable encodings
Use in your systems
<role vocab="credit"
vocab-identifier="https://casrai.org/dictionary/"
vocab-term="Problematic Paper Screener (PPS)"
vocab-term-identifier="https://casrai.org/dictionary/term/problematic-paper-screener" />{
"@context": "https://schema.org",
"@type": "DefinedTerm",
"@id": "https://casrai.org/dictionary/term/problematic-paper-screener",
"name": "Problematic Paper Screener (PPS)",
"identifier": "https://casrai.org/dictionary/term/problematic-paper-screener",
"description": "A free, public automated screening website, built and maintained by Guillaume Cabanac with Cyril Labbé and Alexander Magazinov, that flags candidate papers exhibiting tortured phrases, computer-generated text (SCIgen/Mathgen), or heavy reliance on retracted references ('Feet of Clay'), posting the results to PubPeer for human review rather than issuing findings itself.",
"inDefinedTermSet": "https://casrai.org/dictionary/domain/research-integrity#set",
"url": "https://casrai.org/dictionary/term/problematic-paper-screener",
"sameAs": [],
"license": "https://creativecommons.org/licenses/by/4.0/",
"publisher": {
"@id": "https://casrai.org/#organization"
},
"author": {
"@id": "https://casrai.org/#editorial-team"
},
"datePublished": "2026-08-22T08:45:35",
"dateModified": "2026-09-04T07:25:28",
"inLanguage": "en-GB",
"isAccessibleForFree": true
}







