Skip to main content

From dataflow to insight: Making high-frequency data work for SDTM

A recap of Seiko Yamazaki's P21 webinar on managing high-frequency wearable and device data in SDTM without overwhelming validation, with strategies for design and mapping

Wearables, eDiaries, and continuous monitoring have changed what "clinical trial data" means. They've also created a quiet crisis in SDTM implementation: datasets so large that validation, quality control, and review slow to a crawl.

In this webinar, Seiko Yamazaki walked through why this is happening, showed two real examples of the problem, and laid out practical strategies for keeping high-frequency data usable without sacrificing CDISC compliance.

🪴 Data has outgrown the old assumptions

Clinical data collection has moved well past the visit-based model. eDiaries, wearables, and continuous monitoring technologies now let sponsors capture data more frequently and in real time - a genuine advance, but one that produces much larger and more granular datasets than SDTM was originally built to handle.

The trend is well documented. A 2021 study by Nakazuka found that wearable device use in clinical trials jumped from just one to three studies per year before 2019 to roughly fourteen by 2021.

That growth raises the stakes for CDISC alignment. CDISC standards exist to keep data consistent, standardized, and usable for regulatory review, but the challenge isn't simply standardizing more data. It's structuring fundamentally different data streams (continuous, high-volume, densely time-stamped) in a way that still preserves clinical meaning and traceability for reviewers. That tension, granularity versus usability, is the throughline for everything that follows.

⛓️‍💥 Why high-frequency data breaks the validation model

High-frequency sources can generate thousands, even millions, of data points per subject. Mapped directly into SDTM without planning, domains such as VS or LB balloon in size, and the consequences extend well beyond storage:

  • Validation and quality control processes slow down significantly.

  • Some tools hit performance limits or produce incomplete outputs.

  • Dataset finalization, and ultimately submission timelines, can be delayed.

The root issue is that standard SDTM validation rules were designed for record-by-record checks against conventionally sized datasets. That approach becomes computationally expensive at scale, and it's especially punishing for two categories of checks:

Supplemental qualifiers and relational data. Validating SUPP-- data requires joining subordinate records back to their parent domain, then reconstructing QNAM/QVAL pairs into logical variables before rules can even run. At high-frequency scale, that means tens of millions of additional joins and transformations before validation begins.

Cross-domain and timing checks. Aligning dates and times across DM, AE, EX, and VS, or calculating relative days and sequencing, requires comparisons across datasets - work that scales poorly as record counts climb.

⚠️ Two examples of what goes wrong

Yamazaki shared two real patterns that illustrate how direct, unfiltered conversion of source data into SDTM creates unmanageable datasets.

Example 1: Redundant variables. An XG domain captured glucose values from a monitoring device every fifteen minutes, producing roughly 32 million records. The associated SUPPXG dataset, holding coded values for test type, unit of measure, and result source, reached about 95 million records. The problem: the QLABEL and QVAL values in QNAM were identical across every record for the same variable. Validating this supplemental data means joining back to 95 million largely repetitive rows, adding enormous overhead for very little informational value.

Example 2: Data duplication. A ZA domain captured food intake values alongside primary observations. Its SUPPZA dataset carried event date/time, food name, quantity, measurement, source, and local date/time, much of which duplicated information already present in the parent domain.

Both cases came from treating SDTM as a direct pass-through for source data. The result: validation runs that take an extremely long time, and datasets that aren't actually more clinically meaningful for having preserved every raw value.

✅ Strategy 1: Get the design right before mapping starts

Managing data volume starts well before mapping. A few design-stage principles make the biggest difference:

  • Choose domains deliberately. Domain selection is a real lever; it affects record counts, clarity, validation burden, and reviewer usability. This decision deserves the same rigor as any other part of SDTM planning.

  • Aggregate and summarize when the science supports it. If the real question is about trends, thresholds, or events - not every individual reading - structure SDTM around clinically interpretable summaries instead of raw streams. Yamazaki's example: a wearable capturing heart rate every second generates over 86,000 records per subject per day. But if the safety question is tachycardia occurrence, reviewers don't need every heartbeat. They need each predefined clinically significant episode captured as a record in a clinical events domain. Any aggregation approach needs scientific justification, clear documentation, and a traceable link back to the raw data.

  • Align early, across functions. Standards, programming, and regulatory teams should agree on the level of detail before mapping begins. This isn't a nice-to-have. It's what prevents costly rework, keeps SDTM consistent with analysis datasets, and ensures the data meets regulatory expectations from the start.

The underlying principle: SDTM is not meant to replicate a source database or device output stream. It's a standardized, review-friendly format built for clinical meaning and regulatory usability, and both examples above show what happens when that principle gets lost.

🗺️ Strategy 2: Optimize at the mapping stage

Even with good design, the mapping stage is where complexity either gets controlled or gets locked in. Key tactics:

  • Be intentional about record volume. Look for ways to reduce records intelligently during mapping rather than inheriting the consequences downstream.

  • Limit cross-domain linkages to ones that add real value. Every linkage increases structural complexity and validation workload. Keep what supports review; cut what doesn't.

  • Map only what supports analysis, interpretation, or review. System-generated variables or other unused information generally shouldn't become supplemental qualifiers; they inflate dataset size and validation complexity without adding reviewer value. In Example 1, since the QNAM values were identical across all SUPPXG records, that information could instead live in parent-domain variables like XGSPID or XGREFID.

  • Use SEQ as the ID variable for supplemental data. Linking supplemental records to their parent via the sequence variable creates a precise one-to-one relationship. Using variables like VISIT should generally be avoided, since a single visit can span multiple records and unintentionally link one supplemental record to several parent records. Getting IDVAR/IDVARVAR right also sets up a smoother transition to the new SDTM Implementation Guide's approach to nonstandard qualifiers.

  • Keep datasets focused by concept. Avoid combining device data, observation values, and contextual annotations into one oversized dataset. Different concepts deserve different, purpose-built domains; combining them just makes validation and review harder.

  • Reassess "not done" and "occur = N" records at scale. In extremely large datasets, records that don't represent actual observations can add substantial bulk without adding review value. This isn't a general recommendation, but when non-event records are significant enough to affect dataset size, sponsors should evaluate whether they're necessary, and if so, whether they belong in a separate custom domain rather than the primary domain. Document any decision to omit them and confirm it doesn't affect interpretability or CDISC conventions.

🔧 Strategy 3: When you can't shrink it, adjust the workflow

Sometimes a parent domain and its associated SUPP-- datasets simply can't get smaller. In that case, the fix shifts from data volume to validation strategy:

  1. Validate large parent domains, their sub-core datasets, and DM separately, so subject-level checks can run without waiting on everything else.

  2. Validate the remaining submission package as its own group.

  3. Schedule a final, comprehensive validation of the full dataset overnight or over a weekend, once intermediate issues are resolved.

This staged approach doesn't reduce total computation, but it surfaces issues faster and keeps teams from being blocked on a single marathon validation run.


🗝️ Key takeaways

  • SDTM compliance doesn't require maximal granularity. Not every collected data point needs to appear at its highest level of detail. Datasets should give reviewers enough to understand what happened in the study without complexity that adds no review value.

  • Design for review needs, not collection frequency. Modeling decisions should reflect how reviewers evaluate safety, efficacy, and study conduct, not simply how often or how densely data was captured at the source.

  • Cross-functional alignment early prevents rework later. Standards, programming, and regulatory teams should agree on modeling strategy during study planning, not after mapping is underway.

  • When in doubt, ask the agency. Proactive engagement with regulatory review divisions can clarify expectations, surface concerns early, and reduce the risk of rework or review delays.


⁉️ Q & A

Q: Can you elaborate on why SUPP-- becomes such a bottleneck?

A: Sub-core datasets require relational joins back to the parent domain during validation. With large, high-frequency datasets, that can mean tens of millions of additional joins and transformations before validation rules even execute, and it gets worse when the supplemental variables themselves are repetitive. Minimizing unnecessary SUPP-- usage can dramatically improve performance.

Q: Why is this becoming a bigger issue now?

A: More trials are using wearable devices and remote monitoring technologies, and those tools generate far larger datasets than traditional data collection methods.

Q: Why can't we just put all the raw data into SDTM?

A: Technically, you can, but it tends to create extremely large datasets that are difficult to validate and review, as shown in the glucose monitoring and food intake examples. More data isn't always more useful, especially for regulatory review.

Q: What's the main problem with very large SDTM datasets?

A: Performance. Validation slows significantly, datasets become harder to review, and processing can take a long time.

Q: Should every device data point be submitted?

A: Not necessarily. The priority should be submitting data that's meaningful and useful for reviewers, not every point a device happens to capture.

Q: What's your recommendation for teams working with wearable data?

A: Start planning early. Standards, programming, statistics, and regulatory teams should align on the modeling approach before SDTM mapping begins.

Q: Is this mainly a technical problem or a standards problem?

A: Both. There are real technical challenges in processing large datasets, but there are equally important design decisions in how that data should be modeled.

Q: Do you expect this issue to keep growing?

A: Yes. As wearables and remote monitoring become more common, handling high-frequency data efficiently will only become more important.

Did this answer your question?