
Regulatory Document Knowledge Mining and OCR for an Ophthalmic Device Manufacturer
Let's Connect
Overview
What we built
An ophthalmic device manufacturer held decades of regulatory and quality records mostly on paper and scanned PDFs across 12 sites and markets, so even routine audits meant days of searching filing rooms. We built an OCR and knowledge-mining pipeline that turns that paper archive into a searchable repository.
In plain terms: an ophthalmic device manufacturer operating across 12 sites and markets was sitting on decades of regulatory submissions, quality-system records and test reports, most of it on paper or in scanned PDFs. Preparing for an audit meant days of manual retrieval from filing rooms, every renewal submission was rebuilt from scratch each cycle because nothing from the last one could easily be found again, and the institutional knowledge locked inside those documents was effectively invisible to the regulatory staff who needed it most.
We built an OCR and knowledge-mining pipeline that scans, reads and classifies that archive automatically, then extracts the device models, markets, submission types and key dates buried inside it into a searchable repository with semantic search over the full corpus. Over 480,000 pages of regulatory and QMS records were digitised and classified within 7 months, document retrieval during audits fell from an average of 2 days to under 5 minutes, classification accuracy reached 96% across 14 regulatory document types, and renewal submission preparation time dropped roughly 35% across the 12 markets.
The Problem
Decades of paper records
Operating across 12 sites and markets, the manufacturer had accumulated decades of regulatory submissions, quality-system records and test reports, the paper trail every ophthalmic device programme leaves behind. Most of it lived on paper or in scanned PDFs with no shared index, so finding a specific document meant knowing, or guessing, which filing room and which box it might be in.
Audit preparation exposed the cost most directly. Retrieving the right records meant days of manual searching through filing rooms, often across more than one of the 12 sites, before an audit could even begin. Renewal submissions fared no better: rather than building on what had already been submitted, each cycle rebuilt its supporting documentation from scratch, because nothing from the previous submission could be found quickly enough to reuse.
The deepest cost was invisible day to day. Decades of institutional knowledge, how past submissions had been argued, what test reports had already established, which markets required which documentation, sat inside those paper records but was effectively invisible to the regulatory staff who now needed it. Every new hire started from nothing, and every submission repeated work the archive already contained the answer to.
Decades on paper
Regulatory submissions, quality-system records and test reports spanning decades sat on paper or in scanned PDFs with no shared index across the 12 sites.
Slow audit retrieval
Preparing for an audit meant days of manual retrieval from filing rooms, searching across multiple sites for records that should have been findable in minutes.
Rebuilt renewals
Renewal submissions were rebuilt from scratch each cycle because nothing from the previous submission could be located quickly enough to reuse.
Invisible institutional knowledge
Decades of institutional knowledge inside past submissions and test reports was effectively invisible to current regulatory staff, who had no way to search it.
What it was costing them
Every audit began with days of manual searching instead of preparation, and every renewal cycle repeated work the archive already contained the answer to, because nothing was findable quickly enough to reuse. Regulatory staff operating across 12 sites and markets could not draw on decades of institutional knowledge sitting in paper records, so submissions and audits alike cost more time than the underlying documentation should have required.
The Solution
Searchable OCR knowledge repository
We built an OCR and knowledge-mining pipeline around the idea that a scanned page should become a searchable, structured record rather than an image nobody can query. Bulk scanning intake feeds layout-aware text extraction, which reads the document regardless of its original paper layout, and automated classification sorts each one into its regulatory document type as it comes in.
Getting documents classified was only the first layer. Entity extraction pulls out the device models, markets, submission types and key dates buried inside the text, turning unstructured paper into structured fields that can be searched and filtered directly. Everything lands in a repository with semantic search over the full corpus, so staff can search by meaning, not just exact keywords.
Accuracy improves over time rather than being fixed at launch. A human review loop checks classifier output, corrects mistakes and feeds those corrections back into retraining, so the classifiers covering the manufacturer's 14 regulatory document types keep improving as more of the archive passes through the pipeline.
Key decisions
Bulk scan the archive
Decades of paper records across the 12 sites entered the pipeline through bulk scanning intake, turning physical filing rooms into a digitised source of record.
Extract layout-aware text
Layout-aware text extraction reads each scanned page regardless of its original format, converting images of documents into searchable text.
Classify by document type
Automated classification sorts every document into its regulatory type as it is processed, replacing manual filing with structured, searchable categories.
Extract key entities
Entity extraction pulls device models, markets, submission types and key dates out of the text, giving the repository structured fields to search and filter by.
Close the loop with human review
A human review loop corrects classifier mistakes and retrains the models continuously, improving accuracy across all 14 regulatory document types over time.
Measurable Impact
What changed after launch
The pipeline turned a paper archive into a searchable asset. Over 480,000 pages of regulatory and QMS records were digitised and classified within 7 months, and document retrieval during audits fell from an average of 2 days to under 5 minutes, work that once meant searching filing rooms across the 12 sites.
Quality and speed both improved. Classification accuracy reached 96% across 14 regulatory document types after review-loop tuning, and renewal submission preparation time dropped roughly 35% across the 12 markets, because staff could finally find and reuse what earlier submissions had already established.
Record format
Decades of paper and scanned PDFs
480,000 pages digitised and classified
Audit retrieval
An average of 2 days per request
Under 5 minutes per request
Classification
No automated document classification
96% accuracy across 14 document types
Renewal preparation
Rebuilt from scratch every cycle
Roughly 35% faster across 12 markets
Headline results
Over 480,000 pages of regulatory and QMS records digitised and classified within 7 months
Document retrieval during audits cut from an average of 2 days to under 5 minutes
96% classification accuracy across 14 regulatory document types after review-loop tuning
Renewal submission preparation time reduced by roughly 35% across the 12 markets
Tech & Tools Used
What powered the build
Every tool below earned its place in this engagement. Here is the part each one played.
Python
Provides the glue logic across the pipeline, coordinating scanning intake, text extraction, classification and entity extraction into one automated flow.
AWS Textract
Performs the layout-aware text extraction on scanned pages, turning images of decades-old paper records into machine-readable text.
spaCy
Runs entity extraction over the extracted text, pulling out device models, markets, submission types and key dates for the repository.
Hugging Face Transformers
Powers the automated classification models that sort each document into its regulatory document type, retrained continuously through the review loop.
OpenSearch
Indexes the full corpus for semantic search, letting staff find documents by meaning rather than exact filing-room location.
PostgreSQL
Stores document metadata, classification results and extracted entities, giving the repository a structured record behind every scanned page.
Python (FastAPI)
Serves the repository's search and retrieval functionality, connecting the classified, indexed documents to the staff who need them.
React
Builds the interface staff use to search the repository, review classification results and correct the classifiers through the human review loop.
Docker
Packages the scanning, extraction and classification services as containers, keeping the pipeline consistent as it processed hundreds of thousands of pages.
AWS S3
Stores the scanned originals and digitised outputs, giving every one of the 480,000 pages a durable, retrievable home.
Ready to Build your Medical Devices Business with Knowledge Mining & OCR
Ask Byte
Ask Byte
Typically replies instantly
just Now
Hi! I'm OrganByte's assistant. How can I help you today?
AI-generated content may be incorrect

