
Clinical Trial Data Engineering Platform for a Multi-Site Research Organisation
Let's Connect
Overview
What we built
A contract research organisation running trials across 30 sites had data scattered across electronic data capture exports, lab spreadsheets and instrument files, with statisticians spending days reconciling it before every interim analysis. We built a central data engineering platform that ingests, standardises and validates the data automatically every night.
In plain terms: a contract research organisation running trials across 30 sites was capturing data through a patchwork of electronic data capture exports, lab spreadsheets and instrument files, with nothing tying them together. Statisticians spent days manually cleaning and reconciling datasets before every interim analysis, discrepancies surfaced late in the cycle when they were hardest to fix, and none of it was clean enough for the modelling work sponsors were starting to request. The organisation was, in effect, rebuilding its data from scratch before it could even begin analysis.
We built a central data engineering platform: automated connectors ingest nightly from EDC and lab systems, everything is standardised into CDISC-aligned schemas, and rule-based validation flags discrepancies the moment data loads rather than weeks later. Data from all 30 sites now flows through 12 automated connectors, dataset preparation ahead of interim analyses fell from 10 days to under 48 hours, 95% of data-entry discrepancies are caught automatically before statistical review, and the same curated layer now feeds 3 ML-ready datasets for enrolment forecasting and site-performance models.
The Problem
Fragmented trial data capture
Running trials across 30 sites meant the organisation was, in practice, running 30 slightly different ways of capturing data. Electronic data capture exports, lab spreadsheets and instrument files arrived in whatever shape each source produced them, and nothing in the pipeline reconciled them into a single, trustworthy dataset. Every interim analysis started from a blank page instead of a known-good baseline.
Statisticians absorbed the cost of that fragmentation directly. They spent days manually cleaning and reconciling datasets before every interim analysis, checking exports against lab spreadsheets and instrument files line by line for discrepancies that should have been caught automatically. That work had to happen every single cycle, because nothing upstream improved between analyses.
The discrepancies that mattered most were also the hardest to catch this way: they surfaced late in the cycle, often only once statisticians were deep into an analysis, by which point tracing them back to their source took even longer. And sponsors were starting to ask for modelling work, forecasting enrolment and comparing site performance, that the data simply was not clean enough to support.
Fragmented source formats
Data arrived through a patchwork of electronic data capture exports, lab spreadsheets and instrument files, with nothing standardising them into one usable dataset.
Manual reconciliation
Statisticians spent days manually cleaning and reconciling exports, spreadsheets and instrument files before every interim analysis, and the same manual work repeated in full each cycle.
Late-surfacing discrepancies
Discrepancies surfaced late in the cycle, often deep inside an analysis, by which point tracing them back to their source took even longer to resolve.
Not modelling-ready
Nothing in the pipeline produced data clean enough for the modelling work sponsors were starting to request, such as enrolment forecasting and site comparisons.
What it was costing them
Every interim analysis began with days of manual reconciliation instead of statistics, so statistician time went into cleaning data rather than interpreting it. Discrepancies caught late meant rework deep into each analysis cycle, and because nothing in the pipeline produced modelling-ready data, the organisation could not yet deliver the enrolment forecasting and site-performance work sponsors were starting to request across its 30 sites.
The Solution
Automated AI-ready data pipelines
We built a central data engineering platform to take reconciliation out of statisticians' hands entirely. Automated ingestion connectors pull nightly from EDC and lab systems across all 30 sites, so every dataset starts each morning already assembled rather than waiting for someone to gather it from a patchwork of exports and spreadsheets.
Consistency came from standardising everything into CDISC-aligned schemas the moment it landed, so an interim analysis draws from one dataset shape regardless of which site or instrument produced the underlying data. Rule-based validation runs at load time rather than at analysis time, flagging discrepancies while they are still easy to trace back to their source.
The curated layer we built does double duty. Versioned, validated datasets feed statistical review the way they always did, but the same layer now also produces ML-ready feature tables, giving the organisation a foundation for the enrolment forecasting and site-performance modelling sponsors had started asking for.
Key decisions
Automate nightly ingestion
Connectors now pull data from EDC and lab systems across all 30 sites automatically every night, replacing the manual exports statisticians used to gather themselves.
Standardise into one schema
Every dataset is standardised into CDISC-aligned schemas at load time, so statisticians work from one consistent shape regardless of which site or system produced it.
Validate at load time
Rule-based validation flags discrepancies the moment data loads rather than once an analysis is underway, catching problems while they are still easy to trace.
Version the curated layer
Curated datasets are versioned as they are produced, giving statistical review a stable, traceable dataset instead of a fresh reconciliation each cycle.
Build ML-ready feature tables
The same curated layer now produces feature tables for enrolment forecasting and site-performance models, work the previous pipeline could not support.
Measurable Impact
What changed after launch
The platform changed how a cycle begins. Data from all 30 trial sites is ingested nightly through 12 automated connectors, and dataset preparation ahead of interim analyses fell from 10 days to under 48 hours, freeing statisticians from days of manual reconciliation before they can even start analysing.
Quality moved earlier too. 95% of data-entry discrepancies are now caught by automated validation before statistical review even begins, instead of surfacing late in the cycle. The curated layer also now feeds 3 ML-ready datasets supporting enrolment forecasting and site-performance models, the modelling work sponsors had been asking for.
Data ingestion
Manual exports from EDC and lab systems
12 automated connectors across 30 sites nightly
Dataset preparation
10 days of manual cleaning per cycle
Under 48 hours before interim analyses
Discrepancy detection
Surfaced late, deep into analysis
95% caught automatically before review
Modelling readiness
No data clean enough for modelling
3 ML-ready datasets for forecasting models
Headline results
Data from 30 trial sites ingested nightly through 12 automated connectors, replacing manual exports
Dataset preparation time ahead of interim analyses cut from 10 days to under 48 hours
95% of data-entry discrepancies now caught by automated validation before statistical review
3 ML-ready curated datasets feeding enrolment forecasting and site-performance models
Tech & Tools Used
What powered the build
Every tool below earned its place in this engagement. Here is the part each one played.
Python
Provides the scripting layer behind the ingestion connectors, handling the source-specific logic needed to pull data out of EDC and lab systems each night.
Apache Airflow
Orchestrates the nightly ingestion and standardisation jobs across all 30 sites, sequencing each step from raw export to curated, validated dataset.
dbt
Transforms ingested data into the CDISC-aligned schemas, giving every dataset the same standardised shape regardless of its originating site or system.
PostgreSQL
Stores the operational and curated datasets that feed statistical review, keeping versioned data available for every interim analysis.
Snowflake
Warehouses the larger curated and feature-table data, giving the ML-ready datasets for enrolment forecasting and site-performance models a scalable home.
Great Expectations
Runs the rule-based validation at load time, flagging discrepancies the moment data lands instead of once an analysis is already underway.
Python (FastAPI)
Exposes the curated datasets and feature tables through an API, letting statistical review and modelling work draw from the same validated layer.
AWS S3
Holds versioned raw and curated files, giving every dataset produced by the pipeline a traceable, storable record.
Docker
Packages the ingestion and validation jobs as containers, so the same pipeline behaves consistently regardless of which site's data it is processing.
Metabase
Gives statisticians and site teams a way to review curated dataset status and validation results without waiting on a manual report.
Ready to Build your Clinical Research Business with Data Engineering for AI
Ask Byte
Ask Byte
Typically replies instantly
just Now
Hi! I'm OrganByte's assistant. How can I help you today?
AI-generated content may be incorrect

