Hero
Data & Knowledge AI
Knowledge Mining & OCR
Medical Devices

Regulatory Document Knowledge Mining and OCR for an Ophthalmic Device Manufacturer


Let's Connect

Overview

What we built

An ophthalmic device manufacturer held decades of regulatory and quality records mostly on paper and scanned PDFs across 12 sites and markets, so even routine audits meant days of searching filing rooms. We built an OCR and knowledge-mining pipeline that turns that paper archive into a searchable repository.

In plain terms: an ophthalmic device manufacturer operating across 12 sites and markets was sitting on decades of regulatory submissions, quality-system records and test reports, most of it on paper or in scanned PDFs. Preparing for an audit meant days of manual retrieval from filing rooms, every renewal submission was rebuilt from scratch each cycle because nothing from the last one could easily be found again, and the institutional knowledge locked inside those documents was effectively invisible to the regulatory staff who needed it most.

We built an OCR and knowledge-mining pipeline that scans, reads and classifies that archive automatically, then extracts the device models, markets, submission types and key dates buried inside it into a searchable repository with semantic search over the full corpus. Over 480,000 pages of regulatory and QMS records were digitised and classified within 7 months, document retrieval during audits fell from an average of 2 days to under 5 minutes, classification accuracy reached 96% across 14 regulatory document types, and renewal submission preparation time dropped roughly 35% across the 12 markets.

The Problem

Decades of paper records

Operating across 12 sites and markets, the manufacturer had accumulated decades of regulatory submissions, quality-system records and test reports, the paper trail every ophthalmic device programme leaves behind. Most of it lived on paper or in scanned PDFs with no shared index, so finding a specific document meant knowing, or guessing, which filing room and which box it might be in.

Audit preparation exposed the cost most directly. Retrieving the right records meant days of manual searching through filing rooms, often across more than one of the 12 sites, before an audit could even begin. Renewal submissions fared no better: rather than building on what had already been submitted, each cycle rebuilt its supporting documentation from scratch, because nothing from the previous submission could be found quickly enough to reuse.

The deepest cost was invisible day to day. Decades of institutional knowledge, how past submissions had been argued, what test reports had already established, which markets required which documentation, sat inside those paper records but was effectively invisible to the regulatory staff who now needed it. Every new hire started from nothing, and every submission repeated work the archive already contained the answer to.

Decades on paper

Regulatory submissions, quality-system records and test reports spanning decades sat on paper or in scanned PDFs with no shared index across the 12 sites.

Slow audit retrieval

Preparing for an audit meant days of manual retrieval from filing rooms, searching across multiple sites for records that should have been findable in minutes.

Rebuilt renewals

Renewal submissions were rebuilt from scratch each cycle because nothing from the previous submission could be located quickly enough to reuse.

Invisible institutional knowledge

Decades of institutional knowledge inside past submissions and test reports was effectively invisible to current regulatory staff, who had no way to search it.

What it was costing them

Every audit began with days of manual searching instead of preparation, and every renewal cycle repeated work the archive already contained the answer to, because nothing was findable quickly enough to reuse. Regulatory staff operating across 12 sites and markets could not draw on decades of institutional knowledge sitting in paper records, so submissions and audits alike cost more time than the underlying documentation should have required.

The Solution

Searchable OCR knowledge repository

We built an OCR and knowledge-mining pipeline around the idea that a scanned page should become a searchable, structured record rather than an image nobody can query. Bulk scanning intake feeds layout-aware text extraction, which reads the document regardless of its original paper layout, and automated classification sorts each one into its regulatory document type as it comes in.

Getting documents classified was only the first layer. Entity extraction pulls out the device models, markets, submission types and key dates buried inside the text, turning unstructured paper into structured fields that can be searched and filtered directly. Everything lands in a repository with semantic search over the full corpus, so staff can search by meaning, not just exact keywords.

Accuracy improves over time rather than being fixed at launch. A human review loop checks classifier output, corrects mistakes and feeds those corrections back into retraining, so the classifiers covering the manufacturer's 14 regulatory document types keep improving as more of the archive passes through the pipeline.

Key decisions

01

Bulk scan the archive

Decades of paper records across the 12 sites entered the pipeline through bulk scanning intake, turning physical filing rooms into a digitised source of record.

02

Extract layout-aware text

Layout-aware text extraction reads each scanned page regardless of its original format, converting images of documents into searchable text.

03

Classify by document type

Automated classification sorts every document into its regulatory type as it is processed, replacing manual filing with structured, searchable categories.

04

Extract key entities

Entity extraction pulls device models, markets, submission types and key dates out of the text, giving the repository structured fields to search and filter by.

05

Close the loop with human review

A human review loop corrects classifier mistakes and retrains the models continuously, improving accuracy across all 14 regulatory document types over time.

Measurable Impact

What changed after launch

The pipeline turned a paper archive into a searchable asset. Over 480,000 pages of regulatory and QMS records were digitised and classified within 7 months, and document retrieval during audits fell from an average of 2 days to under 5 minutes, work that once meant searching filing rooms across the 12 sites.

Quality and speed both improved. Classification accuracy reached 96% across 14 regulatory document types after review-loop tuning, and renewal submission preparation time dropped roughly 35% across the 12 markets, because staff could finally find and reuse what earlier submissions had already established.

Record format

Decades of paper and scanned PDFs

480,000 pages digitised and classified

Audit retrieval

An average of 2 days per request

Under 5 minutes per request

Classification

No automated document classification

96% accuracy across 14 document types

Renewal preparation

Rebuilt from scratch every cycle

Roughly 35% faster across 12 markets

Headline results

Over 480,000 pages of regulatory and QMS records digitised and classified within 7 months

Document retrieval during audits cut from an average of 2 days to under 5 minutes

96% classification accuracy across 14 regulatory document types after review-loop tuning

Renewal submission preparation time reduced by roughly 35% across the 12 markets

Tech & Tools Used

What powered the build

Every tool below earned its place in this engagement. Here is the part each one played.

Python logo

Python

Provides the glue logic across the pipeline, coordinating scanning intake, text extraction, classification and entity extraction into one automated flow.

AWS Textract

Performs the layout-aware text extraction on scanned pages, turning images of decades-old paper records into machine-readable text.

spaCy logo

spaCy

Runs entity extraction over the extracted text, pulling out device models, markets, submission types and key dates for the repository.

Hugging Face Transformers logo

Hugging Face Transformers

Powers the automated classification models that sort each document into its regulatory document type, retrained continuously through the review loop.

OpenSearch logo

OpenSearch

Indexes the full corpus for semantic search, letting staff find documents by meaning rather than exact filing-room location.

PostgreSQL logo

PostgreSQL

Stores document metadata, classification results and extracted entities, giving the repository a structured record behind every scanned page.

Python (FastAPI) logo

Python (FastAPI)

Serves the repository's search and retrieval functionality, connecting the classified, indexed documents to the staff who need them.

React logo

React

Builds the interface staff use to search the repository, review classification results and correct the classifiers through the human review loop.

Docker logo

Docker

Packages the scanning, extraction and classification services as containers, keeping the pipeline consistent as it processed hundreds of thousands of pages.

AWS S3

Stores the scanned originals and digitised outputs, giving every one of the 480,000 pages a durable, retrievable home.

Ready to Build your Medical Devices Business with Knowledge Mining & OCR

Ask Byte

Ask Byte

Typically replies instantly

just Now

Hi! I'm OrganByte's assistant. How can I help you today?

AI-generated content may be incorrect


OrganByte

Building innovative software solutions that transform businesses and drive digital success.

© 2026 YourCompany. All rights reserved.