Skip to content

A Practical Guide to ADaM Dataset Creation and Traceability

ADaM dataset creation is one of the most important parts of clinical trial statistical programming. It is also one of the areas where technical programming quality and statistical thinking must come together.
The purpose of ADaM is not simply to transform SDTM into another standardized format. ADaM datasets are built to support the planned statistical analyses, reproduce results, and provide clear traceability from outputs back to the underlying study data.
For sponsors, ADaM quality can have a direct impact on submission readiness. Poorly structured ADaM datasets may still generate tables, but they can create problems during QC, reviewer navigation, traceability review, sensitivity analysis, integrated summaries, and regulatory response work.

Start with the analysis, not the data

The first mistake in ADaM development is starting with SDTM and asking what can be converted. The better question is: what analyses need to be supported?
ADaM development should begin with the protocol, statistical analysis plan, estimand strategy, endpoint definitions, data review needs, and planned outputs. The SAP defines the populations, endpoints, methods, sensitivity analyses, subgroup analyses, visit windows, baseline definitions, censoring rules, and derivation logic that ADaM must support.
Only after those requirements are clear should the team determine which SDTM domains and external data sources are needed.
This analysis-first approach ensures that ADaM datasets are not just technically compliant, but fit for purpose.

ADSL: the foundation of subject-level analysis

ADSL is usually the foundation of the ADaM package. It provides one record per subject and contains subject-level variables used across analyses, such as treatment assignment, key dates, population flags, stratification factors, demographic variables, baseline characteristics, and disposition-related information.
A strong ADSL dataset helps ensure consistency across outputs. If treatment variables, population flags, or key dates are inconsistent in ADSL, the issue can propagate across multiple datasets and TFLs.
Important ADSL considerations include:

  • Clear treatment variables for planned and actual treatment.
  • Consistent population flags aligned with the SAP.
  • Correct randomization and treatment start dates.
  • Baseline characteristics used in subgroup or covariate analyses.
  • Stratification variables used for randomization or analysis.
  • Subject disposition variables relevant to analysis populations.
  • Traceability to SDTM domains and source data.

ADSL should be reviewed carefully by programming, biostatistics, and clinical teams because it supports the analytical logic of the entire package.

BDS and OCCDS: choosing the right structure

Many ADaM datasets follow the Basic Data Structure, commonly used for repeated measures and parameter-based analyses. BDS is often used for efficacy endpoints, lab values, vital signs, pharmacodynamic measures, and other quantitative or categorical endpoints.
BDS is powerful because it provides a consistent way to represent parameters, analysis values, baseline values, changes from baseline, visit structure, analysis flags, and other key variables.
OCCDS, the Occurrence Data Structure, may be appropriate when the analysis is based on occurrences or events rather than repeated measures. Examples may include adverse event analyses, concomitant medication analyses, and event-driven summaries.
Choosing the right structure is not just a standards decision. It affects how easily outputs can be generated, how clearly reviewers can understand the data, and how well the dataset supports sensitivity or subgroup analyses.

Traceability is the core requirement

Traceability is one of the defining principles of ADaM. It allows reviewers to understand how an analysis result was produced and how it connects to the underlying data.
Traceability should exist at multiple levels:

  • From TFL result to ADaM dataset.
  • From ADaM variable to derivation rule.
  • From ADaM derivation to SDTM source variables.
  • From SDTM data to collected source data where relevant.
  • From metadata to datasets and programs.

Traceability does not mean every variable must be a direct copy from SDTM. Many ADaM variables are derived. The important point is that the derivation is documented, reproducible, and understandable.
For example, a responder endpoint may require baseline definition, post-baseline visit selection, handling of missing values, threshold logic, and treatment window rules. A reviewer should be able to see how those decisions were implemented and where the required data came from.

Metadata is not administrative

ADaM metadata is often underestimated. Dataset labels, variable labels, origins, derivations, controlled terminology, comments, and value-level metadata all help explain the package.
Weak metadata can make a strong dataset difficult to review. Strong metadata makes the programming logic more transparent and supports regulatory confidence.
Metadata should be developed alongside the datasets, not after programming is complete. When metadata is retrofitted at the end, inconsistencies are more likely.

ADaM specifications should be living technical documents

ADaM specifications are where analysis requirements, source data, derivation rules, metadata, and programming expectations are brought together.

Good specifications should include:

  • Dataset purpose and structure.
  • Source domains and variables.
  • Derivation rules.
  • Population flag logic.
  • Parameter definitions.
  • Visit window rules.
  • Baseline and change-from-baseline logic.
  • Imputation or censoring rules.
  • Controlled terminology.
  • Traceability notes.
  • QC considerations.

Specifications should be reviewed by both programmers and statisticians. In complex studies, medical input may also be needed to confirm endpoint logic or clinically meaningful definitions.

Common ADaM creation issues

Common issues include:

  • ADSL variables not consistently used across datasets.
  • Population flags that do not match the SAP.
  • Derivations implemented correctly in code but poorly documented.
  • Inconsistent parameter definitions across datasets.
  • Baseline logic that differs from the SAP.
  • Visit windows applied inconsistently.
  • Missing traceability to SDTM.
  • Excessive custom variables that are not explained.
  • Metadata that does not match actual dataset content.
  • Late SAP changes not fully reflected in ADaM specifications.

Many of these issues are preventable through early alignment and independent QC.

The role of independent QC

ADaM QC should go beyond comparing programmed outputs. It should challenge whether the dataset correctly supports the analysis.
This includes reviewing derivations, population flags, analysis values, parameter-level logic, metadata, traceability, and consistency with the SAP. Independent QC may also include double programming, specification review, validation checks, dataset comparisons, and output-level review.
For submission packages, QC should also consider reviewer usability. A technically correct ADaM dataset may still need better documentation if the logic is complex.

How Bioforum Supports ADaM Development?

Bioforum’s statistical programming teams support ADaM dataset creation across clinical trials, integrated summaries, legacy conversions, and submission packages. Our approach connects programming execution with biostatistical strategy, data management quality, and regulatory documentation.
This integrated model helps ensure that ADaM datasets are not developed in isolation. They are aligned with the SAP, traceable to SDTM, supported by metadata, quality-controlled independently, and ready to generate reliable outputs.

Key Takeaways

ADaM dataset creation is not simply a programming task. It is the process of translating the statistical analysis plan into traceable, reproducible, review-ready analysis data.
The strongest ADaM packages are analysis-driven, metadata-rich, quality-controlled, and clearly connected to SDTM and final outputs. For sponsors, that level of traceability can make the difference between a package that merely validates and a package that supports efficient regulatory review.

Need Expert ADaM Programming Support?

Need support creating ADaM datasets, analysis specifications, TFLs, or submission-ready reviewer documentation? Bioforum’s statistical programming experts help sponsors build traceable and analysis-ready packages aligned with CDISC and regulatory expectations.

Learn more about our services