Dictionary domainTrack A
Generative AI use and disclosure
Vocabulary for human–AI collaboration on research outputs and the disclosure required.
For implementers
Operational deployment checklist for Generative AI use and disclosure: prerequisites, five deploy steps, integration notes for Pure, Symplectic Elements, Worktribe, DSpace, and more, plus the pitfalls that recur in the field.
Terms in this domain
51 terms
EU AI Act
Regulation (EU) 2024/1689, the European Union's law regulating AI systems and general-purpose AI (GPAI) models placed on or used in the EU market. It classifies AI systems into four risk tiers (unacceptable, high, limited, minimal) with obligations scaling to risk, applies a separate GPAI-model regime, and phases in on a multi-year timeline set by Article 113 and later revised by the 2026 'Digital Omnibus' amendment.
AI Literacy (EU AI Act Article 4)
Under Article 4 of the EU AI Act, AI literacy is the obligation on providers and deployers of AI systems to take measures ensuring, to their best extent, a sufficient level of AI literacy among staff and other persons who operate or use AI systems on their behalf -- accounting for those individuals' technical knowledge, experience, education, and training, and the context and population the AI system is used on or for. The obligation entered into force on 2 February 2025 under Article 113(a), ahead of the Act's later staged compliance deadlines for high-risk and GPAI provisions.
C2PA Content Provenance
C2PA content provenance refers to the tamper-evident metadata record — called a Content Credential — attached to a piece of digital media (image, video, audio, or document) under the technical specification published by the Coalition for Content Provenance and Authenticity (C2PA). C2PA is a standards initiative founded by Adobe, Microsoft, Intel, the BBC, and Truepic, since joined by a large roster of member organizations including Google, Meta, OpenAI, Amazon, and Sony. Its specification (versioned; C2PA 2.x as of 2026) defines a cryptographically signed, chained record of who or what created a piece of media, what tools or AI models touched it, and what edits were subsequently made — designed so a viewer can inspect that chain without having to trust a single central authority.
AI Output Verification
AI output verification is the practice of checking generative-AI-produced content — text, code, data analysis, citations, or images — for accuracy, validity, and fabricated material before it is used, published, or relied upon in a research context. It is the necessary corollary to disclosure: telling a reader that an AI tool was used does not, on its own, discharge a researcher's responsibility for whether what the tool produced is actually correct. ICMJE's recommendations state that authors are responsible for all aspects of their work, including any content produced with the assistance of an AI tool, and must vouch for its accuracy and for the absence of plagiarism; COPE's position statement on AI tools similarly places responsibility for verifying AI-generated content on the human author rather than the tool, and does not permit an AI tool to be listed as an author precisely because it cannot take on that accountability.
AI Transparency Marker
An AI transparency marker is a visible label, watermark, or embedded machine-readable metadata tag attached to a piece of content — text, image, audio, or video — to disclose that it was generated or substantially modified by an artificial intelligence system. The concept spans two related but distinct mechanisms: human-readable disclosure (a statement in a manuscript, caption, or figure legend that AI assistance was used) and technical marking embedded directly in the file itself (for example, a C2PA Content Credential or an invisible watermark) that software can detect even if the disclosure text is stripped or ignored.
NIH AI Policy
NIH AI policy is not one document but three separate NIH Guide notices, each governing a different actor at a different stage of the grant lifecycle: NOT-OD-23-149 (June 2023) prohibits NIH scientific peer reviewers from using generative AI tools to analyze applications or draft critiques, on confidentiality grounds; NOT-OD-25-132 (July 2025) tells applicants that NIH will not treat an application substantially developed by AI as an original idea, and caps any PI to six new/renewal/resubmission/revision applications per calendar year; and a May 2026 Extramural Nexus notice extends the standard federal research-misconduct definition (fabrication, falsification, plagiarism) to AI-related lapses in the conduct and reporting of already-funded research. Which rule applies to a given situation depends on which lifecycle stage - review, application, or funded-research conduct - it falls into.
NSF AI Dear Colleague Letter (DCL)
An NSF AI Dear Colleague Letter (DCL) is any of several topic-specific Dear Colleague Letters the National Science Foundation has issued that have artificial intelligence as their explicit subject -- typically flagging a funding priority and directing proposers toward an existing mechanism (RAPID, Planning Grant, EAGER, or a standing program) rather than creating an independent competition. To count as one, a document must (a) be formally issued as a DCL under NSF's Proposal & Award Policies & Procedures Guide (PAPPG) Chapter I definition, carrying an NSF publication number in the nsf.gov DCL series, and (b) name AI/generative AI as the letter's subject. This is a narrower category than 'NSF AI policy' generally, and it explicitly excludes NSF's December 2023 guidance on generative-AI use in the merit review process, which was issued as a Notice to the Research Community, not a DCL -- a distinction searchers frequently get wrong and this page exists to clarify.
ICMJE Generative AI Policy
The International Committee of Medical Journal Editors (ICMJE) sets out its rules on generative AI in Section II.A.4 of its Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals, added in the May 2023 update. The policy rests on three linked rules: (1) chatbots and other AI tools cannot be listed as authors because they cannot take responsibility for the accuracy, integrity, and originality of the work and cannot approve a final version for submission, so they fail ICMJE's own authorship criteria by definition; (2) at submission, authors must disclose whether AI-assisted technologies were used, with the disclosure location depending on how the tool was used — writing, editing, or proofreading assistance is described in the Acknowledgments section, while use of AI to collect or analyze data or to generate figures is reported in the Methods section; and (3) human authors remain fully responsible for any submitted material produced with AI assistance, including reviewing and editing AI output carefully, because such tools can generate authoritative-sounding text that is incorrect, incomplete, or biased.
DOE AI Policy
"DOE AI policy" most precisely refers to the U.S. Department of Energy's Generative Artificial Intelligence Policy (DOE P 2031, "Use of Generative Artificial Intelligence," issued December 29, 2025) -- a departmental directive governing how DOE's own federal workforce and National Laboratory staff may adopt, deploy, and use generative AI tools in their work. It sits inside DOE's broader 200-series management directives and was written to align with the October 2023 White House Executive Order on AI and the March 2024 OMB Memorandum M-24-10 on federal agency AI governance. It is an internal governance instrument, not an applicant-facing submission requirement: it does not create a DOE-wide rule that external grant applicants, proposers, or awardees must disclose generative AI use when submitting a funding opportunity announcement (FOA) response. Where DOE's policy does touch research practice directly is inside the agency: when generative AI contributes to an idea, approach, or invention developed at DOE or a National Laboratory, DOE staff are expected to identify that specific contribution and cite the GenAI tool as part of the research methodology, consistent with the department's broader research-integrity and record-keeping expectations. Research administrators should not assume this term is interchangeable with the applicant-facing generative AI disclosure notices published by other funders -- see NSF AI Policy and Nature Portfolio AI Policy for that different category of policy -- and should always check the specific DOE FOA's own instructions and applicable acquisition regulations for any submission-level AI-disclosure requirement, since DOE has not, as of this writing, published a separate proposer-facing AI-disclosure notice comparable to NSF's.
IEEE Generative AI Policy
IEEE has no single unified 'AI policy' document; instead it implements two distinct sets of generative-AI rules, disseminated consistently across its journals, transactions, and conferences via the IEEE Author Center and individual IEEE technical societies. The author-side track requires disclosure of generative-AI use in preparing a manuscript, forbids listing an AI system as a co-author, and holds human authors fully accountable for the accuracy and originality of every word, figure, and citation regardless of how it was drafted. The reviewer-side track is a separate and stricter confidentiality rule: reviewers may not upload any part of a manuscript under review to a public generative-AI platform, and may not use such a platform to draft part or all of a review, because doing so risks exposing confidential, unpublished material to a system that can retain or learn from submitted input.
Nature Portfolio AI Policy
Nature Portfolio's AI policy is the set of editorial rules, published on nature.com and applied across Nature, its sister research and Nature Reviews journals, Scientific Reports, and other Nature Portfolio titles, governing generative AI use in submitted manuscripts. A submission is treated as compliant only if: (1) no AI tool (large language model, chatbot, or image generator) is listed as an author or co-author, since authorship requires accountability that a tool cannot hold; (2) any substantive use of an LLM or other generative AI tool in preparing the manuscript is disclosed in the Methods section (or an equivalent alternative section for article types without one) — minor AI-assisted copyediting of already-written text does not require disclosure; and (3) no AI-generated or AI-substantially-modified image, video, or figure is used, except in the narrow case of a piece specifically about AI itself, reviewed case-by-case by the editor, or artwork sourced through an agency with a contractual chain that makes its provenance legally clear.
NSF AI Policy
NSF AI policy is not a single rule but a set of related National Science Foundation positions on artificial intelligence that a proposer or awardee needs to track separately: (1) guidance, currently framed as encouraged rather than mandatory, on disclosing generative-AI use in preparing a proposal; (2) an explicit statement that fabrication, falsification, or plagiarism committed with the assistance of AI-based tools counts as NSF research misconduct under the Proposal & Award Policies & Procedures Guide (PAPPG); and (3) NSF's own AI research-infrastructure investments, chiefly the National AI Research Resource (NAIRR) and the NSF AI Institutes program, which are funding/access policy rather than proposal-conduct policy. Treat these as three distinct tracks that happen to share the same agency and the same underlying 2023-2025 policy push, not one unified 'AI rule.'
OECD AI Principles
The OECD AI Principles are the first intergovernmental standard on artificial intelligence, adopted by OECD member and partner governments in May 2019 and updated in May 2024 to address general-purpose and generative AI. They consist of five values-based principles for trustworthy AI (inclusive growth/sustainable development/well-being; human rights, democratic values and fairness including privacy; transparency and explainability; robustness, security and safety; and accountability) plus five recommendations for policymakers (invest in AI research and development; foster an inclusive AI-enabling ecosystem; shape an enabling, interoperable governance and policy environment; build human capacity and prepare for labour-market transition; and pursue international co-operation for trustworthy AI). As of the 2024 update, 47 countries and jurisdictions adhere to the Principles, including all OECD members, the European Union, and several non-member states. A research institution or research-performing organization can treat the Principles as an operational baseline for AI governance when: (1) it can point to a documented AI risk-management or oversight process addressing all five values-based principles (not just one, e.g. privacy) as applied to a specific AI use case in the research lifecycle; (2) that process is proportionate to the AI system's stage and context of use rather than a single one-time sign-off; and (3) it is paired with actual transparency to affected parties (participants, authors, reviewers) about where and how AI was used. Simply having an internal 'AI policy' document does not, on its own, satisfy the Principles — the OECD frames them as principles for actors across the AI system lifecycle, not a checklist to file away.
ACM Policy on Authorship
The ACM Policy on Authorship is the Association for Computing Machinery's publisher-wide rule set, applied across its journals, transactions, magazines, and 70+ conference proceedings, defining who may be listed as an author on an ACM-published work, prohibiting generative AI tools from being listed as authors, and requiring authors to disclose any use of generative AI tools and technologies in producing the work — with a narrow exemption for basic word-processing aids such as spelling and grammar checkers.
Fake Citation
A fake citation (also called a fabricated citation or, when the source is a paper never written at all, a phantom reference) is a reference that does not correspond to a real, findable publication -- an invented author, title, journal, DOI, or page range, or some combination of these, presented as if it points to an actual source. It also covers the narrower case of a citation to a real work that is misrepresented: the cited paper exists, but it does not say, show, or support what the citing text claims it does. A citation is operationally 'fake' if a reader who tries to locate and check it cannot verify that the source exists as described, or finds that the source exists but does not support the claim attached to it -- as opposed to an honest error (a typo in a volume number, a wrong year) that still resolves to the intended real work.
Consensus (AI Academic Search Engine)
Consensus is a named AI-powered academic search engine (built by the company Consensus, at consensus.app) that retrieves peer-reviewed papers relevant to a natural-language research question and generates a synthesis of what the retrieved literature says, rather than returning a plain ranked list of results. For questions phrased as a yes/no/maybe claim, it additionally displays a 'Consensus Meter' -- a visual indicator of how the retrieved papers' findings line up (agree, disagree, or mixed) on that specific claim, generated from the paper set Consensus itself retrieved and summarized. It is a specific product in the 'literature summarization and evidence-synthesis' category, distinct from general-purpose AI chatbots (which do not search a dedicated indexed academic corpus or cite retrieved papers by default) and from citation-graph or general-purpose scholarly search tools such as Semantic Scholar (which surface and rank papers but do not generate a claim-level agreement synthesis across them).
AI Research Tool
An AI research tool is software that applies machine learning — typically a large language model (LLM), an embedding-based semantic search index, or both — to a specific stage of the research workflow: finding and screening literature, extracting or summarizing data from papers, mapping citation relationships, drafting or revising manuscript text, or analyzing research data. "AI research tool" is a category label, not the name of any single product: it covers named tools with genuinely different scopes (a citation-mapping tool like Connected Papers does not do what a writing-assistance tool like Paperpal does), and a page or citation that treats it as one interchangeable thing is usually mis-scoped. What makes a tool an instance of this category, rather than a general-purpose AI assistant that happens to get used for research, is that it is built or marketed specifically around a research task — an academic search index, a citation graph, a manuscript-formatting model trained on published literature — rather than being a general chat interface pointed at an arbitrary prompt.
Elicit (AI Research Assistant)
Elicit is a named AI research-assistant platform (built by the public benefit corporation Elicit, spun out of the nonprofit lab Ought in 2023) that searches an indexed academic-literature corpus and performs structured evidence-extraction tasks -- summarizing papers, extracting and tabulating data points across many papers with sentence-level source citations, and supporting systematic-review-style screening and data-extraction workflows aligned to PRISMA 2020. It is a specific product, not a generic label for "AI that helps with research" -- distinct from citation-graph discovery tools (Connected Papers, ResearchRabbit) and general-purpose AI writing assistants -- and its extraction/screening output requires independent verification against the source papers rather than being treated as ground truth.
Model versioning
The practice of identifying a specific revision of an AI model by name, version number, release date, or content hash, sufficient to uniquely distinguish it from earlier or later revisions that may behave differently on the same input.
Inference
The process of generating outputs from a trained AI model in response to inputs at runtime, distinct from training (which updates model parameters); for LLMs, inference is the production of completions from prompts.
Fine-tuning
The process of further training a pre-trained foundation model on a smaller, task-specific or domain-specific dataset, updating some or all parameters, to specialise its behaviour while retaining general capability.
Retrieval-augmented generation (RAG)
An AI architecture in which an LLM is augmented at inference time with documents retrieved from an external corpus (often via vector similarity search), so that the model's outputs are grounded in retrieved evidence rather than relying solely on parametric knowledge.
AI in literature search
The use of AI-powered discovery and retrieval tools (e.g. Elicit, Consensus, Scite, Undermind, Semantic Scholar's semantic-search features, or a general-purpose LLM queried directly) to identify, rank, filter, or synthesise relevant scholarly literature, as distinct from using AI to condense text a researcher has already found.
AI in qualitative coding
The use of an LLM or other AI tool to assign codes, categories, or themes to qualitative data -- interview transcripts, open-ended survey responses, field notes -- either as a first-pass triage that a human coder then reviews, or, more controversially, as a sole or primary coder alongside or instead of a human.
AI image generation
The underlying technical capability of a generative AI system to produce novel images from a text prompt or other input (text-to-image synthesis), as distinct from the separate question of whether a specific generated image is being used appropriately -- disclosed and illustrative, versus undisclosed and presented as a real research result (see synthetic image).
AI summarisation
The use of a generative AI system to condense a longer text -- a paper, dataset documentation, meeting notes, or a set of retrieved sources -- into a shorter form, typically preserving what the tool judges to be the key points while omitting detail, qualification, and nuance.
AI translation
The use of an AI system -- most commonly a neural machine translation model or a general-purpose LLM used for translation -- to convert text from one natural language into another, applied to manuscripts, participant-facing materials, survey instruments, or informed consent documents.
AI in editorial decisions
The use of AI or generative AI tools by a journal editor in the process of managing or deciding on a submitted manuscript -- including drafting decision letters, summarising reviewer reports, or (in the position no major publisher currently permits) using AI output to make or substantially inform the accept/reject/revise judgment itself.
AI in peer review
The use of generative AI by a peer reviewer in the process of evaluating a submitted manuscript, most consequentially by uploading manuscript text (in whole or part) to a third-party AI tool -- an act that most major publishers now explicitly prohibit because the manuscript is confidential, unpublished material.
Generative-AI disclosure statement
A dedicated section in a manuscript — typically headed 'Use of Generative AI' or similar — that consolidates all disclosures of AI tool use across the work, including tools used, versions, sections affected, and the human authors' verification process.
AI co-authorship rejection (ICMJE 2023)
The 2023 update to ICMJE's Recommendations stating explicitly that chatbots and generative AI systems cannot be listed as authors because they cannot satisfy any of the four ICMJE authorship criteria, in particular the requirement to be accountable for the work and to approve the version to be published.
Author responsibility (for AI use)
The principle that human authors retain full responsibility for the accuracy, integrity, originality, ethical sourcing, and lack of plagiarism of all content in a scholarly work, regardless of which portions were drafted, suggested, or generated by an AI tool.
Acknowledgement (vs authorship for AI)
The convention, codified by ICMJE and COPE (2023), that AI tool use must be disclosed in the methods or acknowledgements section of a scholarly work rather than via the author byline or CRediT contributor list, because AI cannot satisfy authorship's accountability requirements.
Training data provenance
A documented record of where an AI model's training data came from -- source, licensing basis, collection method, and chain of custody -- as distinct from what the data actually contains (see training data composition) and from labelling the origin of a specific piece of AI-generated output (see AI provenance / C2PA).
Data leakage (training)
Contamination of an AI model's training corpus with data that should have remained held out for evaluation -- most consequentially, public benchmark questions, answers, or test sets -- which inflates the model's reported performance on that benchmark relative to its true generalisation ability.
AI fairness
A contested normative criterion for evaluating whether an AI system's treatment of different groups or individuals is acceptable -- not a single measurable property, but a family of mutually incompatible formal definitions (e.g. demographic parity, equalised odds, predictive parity/calibration) among which a system generally cannot satisfy more than one simultaneously whenever base rates differ across groups.
AI bias
A measurable, descriptive property of an AI system's outputs: a systematic skew that produces unjustified differences in accuracy, treatment, or representation across groups, tasks, or contexts, arising from training data composition, algorithmic design choices, or deployment context -- distinct from ai-fairness, a contested normative judgment about which disparities are acceptable.
Detection tool (AI-generated)
A software system that estimates the probability that a given piece of content -- typically text, sometimes images -- was produced by a generative AI system, usually by statistically analysing patterns in the content itself without any cooperation from or signal embedded by the original generator, as distinct from watermarking, which requires generator-side cooperation at the point of creation.
AI provenance
The general practice of tracking and documenting the origin of AI-related content across the AI pipeline -- covering both an AI system's inputs (see training data provenance) and the origin of a specific generated output (which model, version, and generation event produced it) -- as distinct from any single concrete technical mechanism for asserting that origin.
Watermarking (AI output)
The embedding of a statistical, cryptographic, or visible signal into AI-generated content at the moment of generation, allowing later identification of that content as AI-produced -- a proactive, generator-side mechanism, as distinct from a detection tool that infers AI origin after the fact without the generator's cooperation.
Synthetic image
An AI-generated or otherwise artificially fabricated image presented within a research context -- most consequentially, an image standing in for genuine experimental or observational data in a publication (e.g. microscopy, blots, clinical imaging) without disclosure, which most journals now treat as a serious integrity breach distinct from disclosed, clearly-labelled illustrative use.
Synthetic data
Data generated artificially -- via simulation, statistical modelling, or a generative AI system -- to mimic the statistical properties of real data without any of its records corresponding to a real individual observation, used for privacy-preserving research, method testing, training-set augmentation, or teaching.
Hallucination
An output from a generative AI system that is presented confidently and fluently but is factually incorrect, fabricated, or unsupported by the input data or any verifiable source — including invented citations, non-existent authors, false statistics, and incorrect quotations.
System prompt
A instruction set or context injected by an AI platform or application developer before any end-user input, configuring a model's persona, constraints, or behaviour for an entire session or deployment -- typically invisible to and not authored by the end user, as distinct from prompt engineering, the user's own iterative practice of crafting their own input.
Prompt engineering
The practice of designing, refining, and structuring the input text (prompt) given to a generative AI system to elicit a specific desired output -- including techniques such as role assignment, few-shot examples, chain-of-thought scaffolding, and output-format specification -- considered here specifically for its reproducibility implications in research work, as distinct from system prompt, the separate, developer-set instruction layer a deploying application configures independently of the end user.
Generative AI
Artificial intelligence systems whose primary output is novel content (text, images, audio, video, code, or structured data) produced by sampling from a learned distribution, as distinct from discriminative AI systems whose output is a classification, score, or decision over existing inputs.
Large language model (LLM)
A neural-network model trained on large text corpora using self-supervised next-token prediction (or analogous objective), with parameter counts typically in the billions, capable of generating coherent text and performing a broad range of natural-language tasks without task-specific training.
AI tool disclosure
A statement within a scholarly work that identifies which generative AI tools were used, the version, the scope of use (e.g., language editing, code generation, figure creation), and which sections were affected, sufficient for a reader to assess the AI's role.
AI as author
The disallowed practice of listing a generative AI system (e.g., ChatGPT, Claude) in the author byline or contributor list of a scholarly work, on the rationale that AI cannot meet authorship criteria requiring accountability, agreement, and the capacity to take public responsibility.
AI-generated content
Text, images, code, or other artefacts produced substantively by a generative AI system in response to a prompt, where the AI is the proximate source of the content rather than a tool refining human-authored material.
AI-assisted writing
The use of a generative AI tool by a human author to draft, edit, paraphrase, summarise, or stylistically revise text where the human retains final editorial control and authorship responsibility.








