Written and maintained by CASRAI Editorial Board
Last updated
The question that decides whether CDISC standards apply to your study is not “is this a drug trial?” It is when the study started, and which application it lands in. FDA’s requirement attaches per study, on a date threshold, and a study that started one day the wrong side of that line is treated completely differently from one that started the day after. Get that wrong and the consequence is not a review comment — it is an automated rejection of the electronic submission before a reviewer ever opens it.
This guide covers what CDISC is, who maintains it, what each standard in the family actually does, and the specific mechanics — date thresholds, validation rule numbers, file constraints — that determine whether a submission is accepted.
What CDISC is and who maintains it
FDA’s own Study Data Technical Conformance Guide describes it plainly: the Clinical Data Interchange Standards Consortium “is an open, multidisciplinary, neutral, nonprofit standards development organization (SDO) that has been working through consensus-based collaborative teams to develop global data standards for clinical and nonclinical research.”
CDISC describes itself as a 501(c)(3) global nonprofit charitable organization with administrative offices in Austin, Texas, governed by a Board of Directors with a Leadership Team and a CDISC Advisory Council. Standards are drafted by volunteer teams drawn from sponsors, CROs, regulators, academia and technology vendors, and go through public review before publication. Controlled terminology — the permitted value sets for coded variables — is developed and published by CDISC in collaboration with the National Cancer Institute’s Enterprise Vocabulary Services (EVS).
The practical consequence of that governance model is worth stating: CDISC is not a regulator. CDISC publishes; FDA decides which published versions it will accept and when. Those are two separate documents maintained by two separate organizations, and conflating them is the root of most version confusion.
The family, and what each member is for
The standards divide cleanly by where they sit in the data lifecycle.
- CDASH (Clinical Data Acquisition Standards Harmonization) governs collection — CRF and eCRF field naming, question text and prompts. The current implementation guide is CDASHIG v2.3, published 28 September 2023, referencing CDASH Model v1.3. CDASH deliberately reuses SDTM domain names so that collection maps onto tabulation with less manual work.
- SDTM (Study Data Tabulation Model) governs tabulation of human clinical trial data. FDA defines it as a standard structure for human clinical trials tabulation datasets. Observations are organised into general classes — Interventions, Events and Findings — plus Findings About, special-purpose domains, and the Trial Design Model that describes the planned conduct of the study.
- SEND (Standard for Exchange of Nonclinical Data) is an implementation of SDTM for animal toxicology studies. SENDIG v3.1.1 (30 March 2021) covers single-dose and repeat-dose toxicology, carcinogenicity and safety pharmacology; SENDIG-DART covers developmental and reproductive toxicology; SENDIG-Genetox v1.0 (28 June 2023) covers genetic toxicology; SENDIG-AR v1.0 covers studies under FDA’s Animal Rule.
- ADaM (Analysis Data Model) governs analysis. FDA is explicit about why it exists separately: SDTM and SEND “do not always provide the data structured in a way that supports all analyses needed for review,” so sponsors should supplement SDTM with ADaM datasets. ADaMIG v1.3 was published 29 November 2021. Note the asymmetry — FDA states that analysis-dataset specifications for nonclinical toxicology studies have not been developed, so there is no ADaM equivalent on the SEND side.
- Define-XML is the metadata standard: the machine-readable data definition file describing datasets, variables, value-level detail, code lists and origins. It builds on CDISC’s ODM-XML. Analysis Results Metadata (ARM) extends it to describe analysis results.
- Controlled terminology supplies the standard values used inside those datasets. External dictionaries such as MedDRA for adverse event coding sit alongside it, and FDA requires the version of any external dictionary to be stated in both the data definition file and the Trial Summary domain.
If you need the SDTM-versus-ADaM boundary drawn in detail — which variables belong where, and why a derived flag does not belong in a tabulation dataset — see the dedicated comparison of CDISC SDTM vs ADaM and the reference entry for the Study Data Tabulation Model.
Model versus implementation guide: the version trap
SDTM and SDTMIG are different documents with different version numbers, and people routinely quote one when they mean the other. The model defines the metadata framework; the implementation guide applies it to a study type. CDISC states there is always a one-to-one relationship between a version of the standard and a version of an implementation guide, though one model version can underpin several IGs — SDTM v2.0 (29 November 2021) pairs with SDTMIG v3.4, SDTM v1.7 (20 November 2018) pairs with SDTMIG v3.3, and SDTM v2.1 (10 June 2024) pairs with the Tobacco Implementation Guide v1.0.
FDA’s Study Data Technical Conformance Guide narrows this further: for the purposes of that guide, “the terms SDTM, ADaM, and SEND apply to versions only listed and supported by FDA in the Catalog.” A version CDISC has published is not thereby usable in a submission. Always check the current Data Standards Catalog rather than assuming the newest CDISC release is acceptable.
How the Data Standards Catalog actually works
The Catalog is a spreadsheet, not prose, and its four date columns carry precise and distinct meanings that sponsors misread constantly. FDA’s own instructions tab defines them:
- Date Support Begins — when sponsors can start using the standard. “Ongoing” means FDA supported it as of the Catalog’s initial release on 13 June 2011.
- Date Support Ends — a study that started before this date may continue using the version; a study starting after it may not.
- Date Requirement Begins — when use becomes mandatory “for any study that starts after” the date. An empty cell means no requirement date has been set.
- Date Requirement Ends — after which the version can no longer be used for studies starting later.
Every one of those is keyed to study start date, not submission date. For clinical studies FDA defines study start date (the SSTDTC parameter in the Trial Summary domain) as the earliest date of informed consent among any subject who enrolled in the study. That single data point governs whether the requirement attaches at all, which is why FDA validates it before anything else.
The statutory basis is section 745A(a) of the Federal Food, Drug, and Cosmetic Act, added by the FDA Safety and Innovation Act, which lets FDA specify electronic submission formats in guidance rather than by rulemaking. The Catalog footnotes then split the requirement by application type: one set of dates “for NDAs, ANDAs, and certain BLAs” and a later set “for certain INDs.”
FDA staff presented the resulting thresholds at a 2023 public meeting as follows: CDER and CBER clinical studies in NDA, BLA and ANDA submissions that started after 17 December 2016; CDER nonclinical studies in NDA, BLA and ANDA submissions after the same date, and in commercial INDs after 17 December 2017; CBER nonclinical studies after 15 March 2023, across all those application types. Treat these as the shape of the rule and confirm current dates against the live Catalog, which FDA revises periodically.
The Technical Rejection Criteria: where submissions actually fail
Since 15 September 2021 FDA has enforced a subset of eCTD validation rules known as the Technical Rejection Criteria for Study Data. These run automatically at receipt. A high-severity failure means the submission is rejected — not queried, not reviewed.
- 1734 — a
ts.xptdataset carrying study start date must be present for each study in the applicable eCTD sections. - 1735 — correct study tagging file (STF) file-tags must be used for all standardized datasets and their define.xml files.
- 1736 — a demographics (DM) dataset and a define.xml must be present: in Module 4 for SEND data, in Module 5 for SDTM data.
- 1737 and 1738 — added effective 16 March 2023 as medium-severity checks; 1738 validates that STUDYID matches between
ts.xptand the STF.
The applicable sections are specific. Nonclinical: eCTD 4.2.3.1, 4.2.3.2 and 4.2.3.4. Clinical: 5.3.1.1, 5.3.1.2, 5.3.3.1–5.3.3.4, 5.3.4, 5.3.5.1 and 5.3.5.2. FDA noted that the rejection criteria are not applied to clinical studies in commercial INDs. If you are unsure which sections your data lands in, the eCTD structure for clinical submissions maps them.
The failure pattern is remarkably concentrated. Across September 2021 to February 2023, FDA reported 453 IND, NDA and BLA nonclinical studies failing rule 1734; 73% (329 of 453) failed simply because ts.xpt was missing, and 71% were repeat-dose toxicology studies. A further 11% failed for a missing study start date value and 16% for a study ID mismatch. In other words, the dominant cause of rejection was not bad science or poor mapping — it was an absent one-row file.
The simplified ts.xpt
This is the fix most teams miss. When a study predates the applicable threshold, or SEND is not required for it, you still have to tell FDA that — by submitting a simplified ts.xpt containing at minimum STUDYID, TSPARMCD, and TSVAL and/or TSVALNF, with one row of information. Where no study start date is available, TSVALNF is populated with “NA”. A study that is legitimately exempt but ships no ts.xpt is indistinguishable, to the validator, from one that is non-compliant. FDA publishes a Technical Rejection Criteria Self-Check Worksheet so sponsors can run the checks before submitting.
File format mechanics that break submissions
Study datasets are still transported as SAS XPORT Version 5 files (.xpt), a format whose constraints shape everything downstream. From the Technical Conformance Guide:
- Each dataset in a single transport file. Datasets over 5 GB should be split into pieces no larger than 5 GB, placed in a
splitsubfolder, with both the split and non-split versions submitted and the splitting logic explained in the reviewer’s guide. - Variable names: maximum 8 characters. Variable descriptive labels and dataset labels: maximum 40 characters.
- Character column lengths should be set to the maximum length actually used across the study’s datasets — if USUBJID never exceeds 18 characters, set it to 18, not 200. This materially reduces file size.
- ASCII only for names and labels. Avoid unbalanced apostrophes (“Parkinson’s”), unbalanced quotation marks, and unbalanced parentheses, braces or brackets in labels — these are named explicitly as prohibited.
- A separate define.xml per dataset type per study: one for the study’s SDTM datasets, one for its SEND datasets, one for its ADaM datasets. FDA calls the data definition file “arguably the most important part of the electronic dataset submission for regulatory review,” and names insufficient documentation of it as a common reviewer-noted deficiency. Define-XML version 2.0 or later is strongly preferred; a printable define.pdf should accompany it if the define.xml cannot be printed.
- Folder structure is prescribed:
m4for nonclinical andm5for clinical, thendatasets, then a study-level folder, thentabulations/sdtm,analysis/adam,legacy,miscandprofilesas applicable.
FDA also publishes two validation rule sets against these data: FDA Business Rules v1.5 (May 2019), applying to SDTM-formatted clinical studies and SEND-formatted nonclinical studies, and FDA Validator Rules v1.6 (December 2022).
Traceability, reviewer’s guides, and legacy data
FDA states that establishing traceability “is one of the most problematic issues associated with any data conversion,” and that if a reviewer cannot trace study data from collection through to analysis, the review “may be compromised.” Traceability here means an unbroken chain: analysis results in the study report → ADaM analysis datasets → SDTM tabulation datasets → source and CRF data.
The documents that carry the human explanation are the reviewer’s guides — the clinical Study Data Reviewer’s Guide (cSDRG), the nonclinical equivalent, and the Analysis Data Reviewer’s Guide (ADRG). These are where you explain deviations, splitting decisions, unit conversions, and any point where an implementation guide gave no instruction and you chose a solution. FDA’s advice on that last case is to discuss the proposed solution with the review division and document it in the SDRG at submission.
Legacy data — study data in a non-standardised format never listed in the Catalog — must be converted in a way that accounts for traceability. Where a collected element cannot be represented as a standardised element, the reviewer’s guide must explain why, and the legacy data (annotated CRF, legacy tabulations, legacy analysis datasets) may need to be submitted alongside the converted data.
Plan it before the first subject consents
The single highest-leverage step is upstream of all of this. FDA asks sponsors to file a Study Data Standardization Plan (SDSP) early in development, typically under the IND, describing how standardized study data will be submitted. For INDs, NDAs and BLAs it belongs in eCTD section 1.13.9 (General Investigational Plan) or 1.20 (General Investigational Plan for Initial IND). No FDA template is mandated; PhUSE publishes an example. The SDSP should be updated as the programme expands, provided at pre-NDA and pre-BLA meetings, and — for clinical studies going to CBER — the SDSP appendix should reach the review office no later than the End-of-Phase 2 meeting. The cover letter accompanying a data submission should describe how far the latest SDSP was executed.
FDA also recommends implementing SDTM before the study is conducted rather than retrofitting, and notes that traceability is enhanced when data is collected prospectively on standardized CRFs such as CDASH. That upstream discipline is ordinary clinical data management practice, and it is far cheaper than remediation.
Where the transport format is heading
SAS XPORT v5 is a decades-old format and its limits (8-character names, 200-byte character values, no native compression) are widely felt. CDER and CBER, working with CDISC and PhUSE, have conducted preliminary testing of CDISC’s Dataset-JSON exchange standard. FDA’s public position is measured: “Initial results indicate potential use as a replacement for XPT v5,” and the centers will conduct further testing before communicating results. Dataset-JSON is not currently a substitute for XPT in a submission, and any adoption would show up in the Data Standards Catalog first.
Beyond FDA
CDISC states that Japan’s Pharmaceuticals and Medical Devices Agency (PMDA) also requires SDTM, ADaM, Define-XML and Analysis Results Metadata for regulatory submissions. The specific version lists, transition arrangements and technical conformance expectations differ from FDA’s, so a submission strategy that satisfies one agency should not be assumed to satisfy the other — check each agency’s own catalogue or notification.
Frequently asked questions
Is CDISC legally required, or just recommended?
It is required, but indirectly and conditionally. Section 745A(a) of the FD&C Act authorises FDA to specify electronic submission formats in guidance; the Data Standards Catalog then names specific standards and versions and gives a “Date Requirement Begins” for each. Where that date is populated and your study started after it, use of the named standard is mandatory for the covered application types. Where the column is empty, FDA supports the standard but has not required it.
Do CDISC standards apply to investigator-initiated academic trials?
Only if the data will be submitted to FDA in a covered application. The requirement attaches to submission types — NDAs, ANDAs, certain BLAs and certain INDs — not to the trial’s funding source or sponsor type. An academic trial whose data is destined for a marketing application falls in scope; one that is never submitted does not. Many institutions nevertheless adopt SDTM-aligned structures voluntarily for reuse and pooling.
What is the difference between SDTM and SDTMIG?
SDTM is the model — the metadata framework defining observation classes and variable roles. SDTMIG is the implementation guide that applies the model to a study type and specifies actual domains and variables. They version independently but pair one-to-one, and FDA lists both a supported version and a supported implementation guide version in the Catalog.
Why does my submission need a ts.xpt when the study is exempt from CDISC?
Because the validator cannot distinguish “exempt” from “non-compliant” without being told. The simplified ts.xpt is how you assert the study start date — or, using TSVALNF, that no start date applies. FDA data showed a missing ts.xpt was by far the largest single cause of rule 1734 rejections.
Which CDISC version should we use for a study starting now?
The current FDA Data Standards Catalog is the only correct answer, and it changes. Do not select a version from a CDISC release announcement, a vendor datasheet, or a previous programme’s submission. Check the Catalog’s supported and required version rows for your standard, centre and application type at the point the study starts, and record the decision in the SDSP.
Do I need ADaM if I already have SDTM?
For clinical studies, yes — FDA states that tabulation data does not always support the analyses needed for review and that sponsors should supplement SDTM with ADaM. For nonclinical toxicology studies there is no ADaM counterpart; FDA notes that analysis-dataset specifications for those studies have not been developed.
Are CDISC standards free to use?
The published standards are available from cdisc.org. CDISC is a membership-funded 501(c)(3) nonprofit, and membership brings additional access and participation rights — but the requirement in a submission is conformance to the standard, not membership of the organisation.
References
- FDA. Study Data Standards Resources — guidance index, Business Rules v1.5, Validator Rules v1.6, and the Dataset-JSON testing statement.
- FDA. Study Data Technical Conformance Guide (Technical Specifications Document, June 2026) — dataset size and naming limits, define.xml requirements, SDSP, reviewer’s guides, traceability, folder structure.
- FDA. Data Standards Catalog — supported and required standards, versions, and the support/requirement date columns.
- FDA. Study Data Technical Rejection Criteria (SEND Face-to-Face public meeting, April 2023) — rules 1734–1738, enforcement dates, rejection statistics, simplified ts.xpt.
- CDISC. SDTM, ADaM, SEND, CDASH and Define-XML standard pages — versions, publication dates, and scope.
- CDISC. About CDISC — nonprofit status, governance, and standards development model.








