Clinical research data management is the discipline of planning, collecting, cleaning, coding, locking, and preserving the data of a clinical study so that its results can be trusted. It is how a protocol — a document describing what a study intends to learn — becomes a dataset that can actually support that learning: analyzable, traceable, and defensible under regulatory and scientific scrutiny.

The discipline has an established professional literature. The Society for Clinical Data Management's Good Clinical Data Management Practices (GCDMP) describes accepted practice across the full span of the work, from case report form design through database lock and archiving [1]. The ICH E6(R3) Good Clinical Practice guideline — the international standard for trial conduct — makes sponsors responsible for data governance and for ensuring data integrity, traceability, and security throughout the data lifecycle [2], and ICH E8(R1) frames quality as something designed into a study rather than inspected in afterward [3]. The job titles vary — clinical data manager, clinical data associate, research data coordinator — but the underlying work is recognizably the same.

What clinical data managers actually do

The daily work usually includes most of the following, though rarely in a tidy sequence:

  • Turning the protocol into data specifications. Reading the protocol and schedule of assessments and deciding exactly what data points must be collected, in what structure, at which visits (protocol and study design, eCRF design).
  • Designing case report forms. Building the electronic case report forms (eCRFs) that sites and participants will complete, ideally drawing on standard libraries rather than starting from scratch (form and study libraries).
  • Writing edit checks and validation rules. Specifying the automated logic that catches out-of-range values, impossible dates, and missing required fields at the moment of entry (edit checks, data validation).
  • Building and testing the study database. Configuring the system, testing forms and checks against realistic scenarios, and documenting that testing before the first participant is enrolled (validation documentation).
  • Overseeing data capture. Monitoring incoming data from sites, labs, participant-reported instruments, and devices (electronic data capture, laboratory data, participant-reported outcomes).
  • Reviewing data and managing queries. Identifying discrepancies, issuing queries to sites, and tracking them to resolution (query management, source data review).
  • Medical coding. Mapping reported adverse events and medications to standard dictionaries — MedDRA for events and diagnoses, WHODrug for medications (MedDRA coding, WHODrug coding, coding workflow and autocoding).
  • Reconciling data across systems. Confirming that serious adverse events in the study database match the safety database, and that external data (labs, imaging, devices) matches enrollment records (data reconciliation, safety reconciliation).
  • Locking the database. Confirming that queries are resolved, coding is complete, and reconciliation is done — then freezing the data so analysis proceeds from a fixed, documented dataset (database lock).
  • Documenting and preserving everything. Maintaining the audit trail, data dictionaries, and records that let someone reconstruct — years later — what was collected, what changed, and why (audit trails, metadata and data dictionaries, archiving and retention).

The data management lifecycle

The clinical data management lifecycle
  1. Design
  2. Build & Test
  3. Capture
  4. Review & Query
  5. Code
  6. Reconcile
  7. Lock
  8. Deliver & Archive

A conceptual educational model — not a formal universal standard. Real studies adapt, reorder, and repeat these stages, and amendments send them back to the start.

The lifecycle runs alongside the study itself: while investigators are recruiting, treating, and following participants, data managers are capturing, querying, coding, and reconciling what those activities produce. We explore how the two tracks interlock in The Two Lifecycles of Clinical Research, and walk the data track end to end in From Protocol to Database Lock.

The work varies enormously

There is no single correct data management workflow. The same discipline looks very different depending on the study.

Phase I versus phase III. An early-phase study may enroll a few dozen participants at one site, with intensive sampling, rapidly evolving procedures, and a small team handling data hands-on. A phase III trial may span hundreds of sites and thousands of participants across countries, with formal data management plans, dedicated coding teams, blinded data review meetings, and interim locks. Scale changes the tooling, the process formality, and the size of the team — not the underlying logic.

Academic versus industry. Industry-sponsored trials are typically run against submission-grade expectations from day one, often with a CRO performing data management under contract. Academic studies are usually grant-funded, staffed by coordinators who wear several hats, and frequently built on institutional platforms such as REDCap, which was designed specifically for this environment [5]. Academic investigators also carry obligations their industry counterparts delegate to sponsors — including, for NIH-funded work, the data management and sharing plans required under the NIH policy in effect since January 2023 [4]. We compare the two worlds in detail in Academic vs. Commercial Clinical Research.

Interventional versus observational and registry studies. A randomized interventional trial revolves around protocol visits, randomization, and safety reporting. Observational studies and registries often trade that intensity for duration: fewer forms per visit, but longitudinal follow-up over years, heavier reliance on EHR integration and routinely collected data, and a different cleaning posture — you cannot query a health record the way you query a site.

What clinical research data management is not

"Data management" is a heavily overloaded term, and readers arrive here from several adjacent fields. This publication is not about:

  • Enterprise or master data management — the corporate IT discipline of governing customer, product, and reference data across business systems.
  • Database administration — installing, tuning, and backing up database servers. Clinical data managers work in databases; DBAs keep the engines running.
  • Library-science research data management — the institutional practice of curating, describing, and sharing research datasets, often based in university libraries. This field genuinely overlaps with ours — NIH data sharing plans sit at the junction [4] — but its center of gravity is preservation and reuse after the research, not the conduct of a study.

Those are legitimate disciplines with their own literatures; they simply share a name. Our short definition page, What Is Clinical Research Data Management?, makes the same distinction for readers arriving from search.

Where technology fits

Nothing in this guide strictly requires software — clinical data management predates it, and paper CRFs with double data entry ran trials for decades. But every activity above produces records someone must capture, check, reconcile, and preserve, under audit-trail and traceability expectations [1][2] that are difficult to meet without purpose-built systems. That is why nearly all of this work now runs through electronic data capture and related platforms.

The technology landscape is broader than a single product category. We classify systems into six primary-function categories — Study-Specific EDC, Institutional Research Platform, Clinical Trial Operations Suite, Participant Data Collection, Research Data Platform, and Configurable Application Platform — described in our technology landscape overview. When you are ready to evaluate systems, our capability reference explains each function in plain language, the technology directory shows what our research has verified about specific products, and Choosing an Electronic Data Capture System walks through the selection process itself.