
Data Strategy and Governance Programme for a Precision Diagnostics Network
Let's Connect
Overview
What we built
A precision-diagnostics network of 14 labs and clinics could not match a patient's clinical record to their molecular results, and cross-site data sharing was mostly refused rather than governed. We gave the network one data dictionary, one patient-matching approach and one governance model.
In plain terms: across this network's 14 labs and clinics, clinical records lived in separate per-site systems while molecular results sat in silos tied to individual instruments, with no common way to say that a clinical record and a molecular result belonged to the same patient. Internal research requests took weeks of manual assembly to answer, field definitions differed from site to site, and compliance concerns meant most cross-site data sharing was simply refused rather than properly governed.
We delivered a data strategy and governance programme built around a network-wide clinical-molecular data dictionary, a master patient index approach for matching records across sites, tiered access controls with standardised de-identification rules, and a stewardship council with named owners for each data domain. A phased consolidation roadmap then moved sites onto the shared model, with automated quality checks gating every migration. All 14 labs and clinics aligned on the dictionary, which now covers 300+ fields, average turnaround on internal research requests fell from 3 weeks to 4 days, 92% of legacy records matched to the new patient index in the first consolidation phase, and the network's first compliance audit after the programme returned zero data-sharing findings.
The Problem
Clinical and molecular silos
The network had grown to 14 labs and clinics, each running its own systems for clinical records, while molecular results were captured separately at the instrument level in their own silos. Nothing tied the two together: without a common patient identifier, a clinician or researcher could not reliably say that a given clinical record and a given molecular result described the same person.
That gap made research slow and manual. Any internal research request that needed data from more than one site took weeks of manual assembly, as staff hunted down records, matched them by hand and reconciled field definitions that differed between sites. What one lab called a result field, another lab might define differently, so even matched records needed further checking before anyone could trust them.
Compliance made the situation worse rather than better. Faced with uncertainty about how cross-site data sharing should be governed, the network's default was to refuse it, so requests that should have taken days were either abandoned or pushed through a slow manual process, and legitimate research was blocked as often as any genuine risk was avoided.
No common patient identifier
Clinical records and molecular results could not be reliably tied to the same patient across the network's 14 labs and clinics, undermining any cross-site analysis.
Manual research requests
Internal research requests needing data from more than one site took weeks of manual assembly rather than a defined process anyone could rely on.
Inconsistent field definitions
Field definitions differed between sites, so even successfully matched clinical and molecular records needed manual reconciliation before they could be trusted.
Sharing refused by default
Compliance concerns meant most cross-site data sharing was simply refused rather than governed, blocking legitimate research along with any genuine risk.
What it was costing them
Every research request that took weeks of manual assembly was a request that arrived too late to be useful, and every refusal made in place of governance was a study that could not happen at all. Meanwhile clinical and molecular data for the same patients sat unmatched across the network's 14 labs and clinics, so the underlying asset, a complete patient record, existed only in theory until someone was willing to spend the weeks to piece it together.
The Solution
Unified governance and access model
We built the programme around a single source of truth: a network-wide clinical-molecular data dictionary that gave every one of the 14 labs and clinics the same field definitions for the first time. Alongside it, a master patient index approach gave the network a consistent way to match clinical and molecular records to the same patient across sites.
Access needed governing as much as data needed matching, so we added tiered access controls with standardised de-identification rules, and a stewardship council with named owners for each data domain, so decisions about sharing had an accountable owner instead of defaulting to refusal.
Moving 14 sites onto the shared model could not happen at once, so a phased consolidation roadmap sequenced the migration site by site, with automated quality checks gating each step before it counted as complete. That combination let the network trust its own data as it consolidated, rather than migrating first and discovering problems afterward.
Key decisions
One dictionary for all sites
The network-wide clinical-molecular data dictionary, covering 300+ fields, gave all 14 labs and clinics the same field definitions instead of 14 separate ones.
Match patients before sharing data
The master patient index approach was built before access rules were finalised, so governance could be applied to correctly matched records rather than guesswork.
Named owners, not default refusal
The stewardship council assigns a named owner to each data domain, replacing the default of refusing cross-site sharing with an accountable decision.
Gate migrations with quality checks
Automated quality checks gate each site's migration onto the shared model, so consolidation could not outpace the network's confidence in its own data.
Phase the rollout
A phased consolidation roadmap moved sites onto the shared model in sequence, rather than attempting all 14 at once.
Measurable Impact
What changed after launch
The programme replaced weeks of manual assembly with a network the 14 labs and clinics could trust. Average turnaround on internal research data requests fell from 3 weeks to 4 days, and 92% of legacy records matched to the new master patient index in the first consolidation phase, evidence that the data dictionary and matching approach were working as designed.
Governance held up under scrutiny too: the network's first compliance audit after the programme returned zero data-sharing findings, a result the previous approach of refusing most requests could never have earned, since refusal was never actually reviewed. The clinical-molecular data dictionary, now covering 300+ fields, and the stewardship council's named owners give the network a way to keep that record clean as it grows.
Patient matching
No common identifier tied clinical and molecular records
92% of legacy records matched to the master patient index
Research turnaround
Requests took weeks of manual assembly
Average turnaround cut from 3 weeks to 4 days
Data sharing
Cross-site sharing simply refused rather than governed
Zero data-sharing findings in the first compliance audit
Field definitions
Differed between all 14 labs and clinics
One data dictionary covering 300+ fields network-wide
Headline results
14 labs and clinics aligned on a single clinical-molecular data dictionary covering 300+ fields
Average turnaround on internal research data requests cut from 3 weeks to 4 days
92% of legacy records matched to the master patient index in the first consolidation phase
Zero data-sharing findings in the network's first compliance audit after the programme
Tech & Tools Used
What powered the build
Every tool below earned its place in this engagement. Here is the part each one played.
Python
Built the matching logic behind the master patient index, linking clinical and molecular records across the network's 14 labs and clinics.
PostgreSQL
Held the master patient index and its matching results, giving every site one authoritative place to check a patient match.
Snowflake
Hosts the consolidated clinical-molecular data once sites migrate onto the shared model, giving research requests one place to query across the network.
dbt
Modelled the shared clinical-molecular data dictionary, so every migrated site's data follows the same governed field definitions.
Apache Airflow
Orchestrates the phased consolidation, running each site's migration and quality checks in sequence.
Great Expectations
Ran the automated quality checks that gated each site's migration onto the shared model before it counted as complete.
OpenMetadata
Tracks the data dictionary's 300+ fields and their definitions, giving every site a shared reference instead of separate local ones.
AWS S3 + Lake Formation
Stores consolidated records with the tiered access controls and de-identification rules the stewardship council defined.
Docker
Packaged the migration and quality-check tooling so it ran the same way at every one of the network's 14 sites.
Ready to Build your Healthcare & Diagnostics Business with Data Strategy & Governance
Ask Byte
Ask Byte
Typically replies instantly
just Now
Hi! I'm OrganByte's assistant. How can I help you today?
AI-generated content may be incorrect

