Hero
AI Solutions & Engineering
MLOps & Model Deployment
Medical Devices

MLOps Pipeline for Shipping Guidance Models to Handheld Ultrasound Devices


Let's Connect

Overview

What we built

A handheld ultrasound maker had guidance models ready months before clinicians could use them, because every update had to hitch a ride on a quarterly firmware release. We built the pipeline that ships models on their own.

In plain terms: the company's handheld scanners use AI to help clinicians capture good images, and improving that AI means updating the model inside each device. But models could only travel inside full firmware builds, released quarterly and installed unevenly across roughly 50 clinic networks. Finished improvements sat on the shelf for months, nobody could say which device ran which version, and if a release misbehaved the only remedy was a support call and a long wait for the next firmware.

We gave the models their own delivery route. A pipeline now takes each new guidance model from the data science team, checks it automatically on the device hardware, and rolls it out over the air in careful stages, no firmware release required. Updates arrive fortnightly instead of quarterly, 92% of devices run the latest model within 14 days of release, and a faulty release can be rolled back automatically in under 15 minutes instead of triggering a multi-week firmware respin.

The Problem

Model updates trapped in firmware

The device maker's capture-guidance models were the product's edge, and its data science team kept making them better. The delivery mechanism undid that work. Every model change had to be hand-packaged into a full firmware build, and firmware shipped quarterly, so a model finished in the first week of a quarter waited out the rest of it before any clinician benefited.

Distribution was as uneven as it was slow. Across roughly 50 clinic networks, devices updated whenever local processes allowed, and no telemetry reported back which version each device was actually running. The fleet became a patchwork: support handled tickets without knowing which model produced the behaviour being reported, and the data science team could not tell how widely any improvement had actually landed.

The riskiest moments had the fewest options. A bad build could not be rolled back; it could only be endured, escalated through support calls and fixed in a future firmware respin measured in weeks. Meanwhile the quality team assembled release documentation by hand for every release, a slow, repetitive effort that made frequent releases even less attractive.

Models trapped in firmware

Guidance models could only ship inside full quarterly firmware builds, so finished improvements waited months before reaching a single clinician's device.

Unknown fleet state

No telemetry recorded which devices ran which model version across roughly 50 clinic networks, leaving support and data science guessing.

No rollback path

A faulty release could not be reversed, only endured: recovery meant support calls and a firmware respin measured in weeks, not minutes.

Manual release evidence

The quality team compiled release documentation by hand for every release, a repetitive effort that made shipping more often feel impossible.

What it was costing them

The gap between model-ready and model-live was costing the product its pace. Clinicians scanned with guidance months older than it needed to be, the data science team watched finished work idle in a queue, support burned time diagnosing issues without knowing the model version involved, and every faulty build turned into weeks of firmware recovery and hand-written documentation.

The Solution

Over-the-air model delivery pipeline

We engineered an end-to-end MLOps pipeline around the client's existing guidance models, starting from a versioned model registry that gives every model a single, authoritative identity from training through to the field. From there, each candidate release is automatically converted for the edge runtime and validated on the device hardware itself before it is allowed anywhere near a clinic.

Delivery is now decoupled from firmware entirely. Models roll out over the air in stages, reaching a small slice of the fleet first and widening only as telemetry stays healthy. Fleet-wide version telemetry reports exactly which device runs which model, and if anything misbehaves, one-click rollback returns affected devices to the previous version without a firmware change.

Because the pipeline records everything it does, release evidence for the quality team is generated automatically from its audit trail rather than assembled by hand. Documentation stopped being a brake on cadence and became a by-product of shipping, which is what allowed the release rhythm to move from one quarterly firmware bundle to fortnightly over-the-air updates without adding quality workload.

Key decisions

01

A registry as single source

Every guidance model is versioned in one registry, so training, validation, rollout and telemetry all refer to the same authoritative artefact.

02

Validate on the device

Each converted model runs automated validation on the edge runtime and hardware before release, catching conversion and performance problems while they are still cheap.

03

Decouple models from firmware

Over-the-air model delivery runs on its own track, so a model update no longer waits for, or puts at risk, a full firmware release.

04

Stage every rollout

Releases widen through the fleet in steps, with telemetry gating each expansion, so a problem surfaces on a few devices rather than across roughly 50 clinic networks.

05

Make rollback one click

Recovery was designed before speed: any release can be reversed automatically, which is what makes fortnightly shipping safe rather than reckless.

Measurable Impact

What changed after launch

The release rhythm changed first. Model updates now ship fortnightly over the air instead of once a quarter inside a firmware bundle, and they land: 92% of fleet devices run the latest guidance model within 14 days of release, where a full quarter used to leave adoption at roughly 40%. The data science team's improvements now reach clinicians while they are still fresh.

The failure story changed just as much. A faulty release now ends in an automated rollback in under 15 minutes rather than a multi-week firmware respin, and fleet telemetry means support finally knows which model version sits behind any ticket. Release documentation effort for the quality team fell from about 3 weeks to 4 days per release, generated from the pipeline's own audit trail.

Release cadence

One quarterly firmware bundle carrying model updates

Fortnightly over-the-air model releases, independent of firmware

Fleet adoption

Roughly 40% on the latest model after a quarter

92% of devices updated within 14 days of release

Faulty-release recovery

Multi-week firmware respin driven by support calls

Automated rollback completed in under 15 minutes

Release documentation

About 3 weeks of manual effort per release

4 days, generated from the pipeline's audit trail

Headline results

Model release cadence improved from one quarterly firmware bundle to fortnightly over-the-air updates

92% of fleet devices ran the latest guidance model within 14 days of release, up from roughly 40% after a full quarter

Faulty-release recovery cut from a multi-week firmware respin to an automated rollback in under 15 minutes

Release documentation effort for the quality team reduced from about 3 weeks to 4 days per release

Tech & Tools Used

What powered the build

Every tool below earned its place in this engagement. Here is the part each one played.

Python logo

Python

The pipeline's working language: packaging, conversion and validation steps, along with the services that manage rollouts and telemetry, are all written in it.

PyTorch logo

PyTorch

The framework the client's guidance models are trained in, and the starting format for every model entering the release pipeline.

ONNX Runtime logo

ONNX Runtime

The edge runtime on the handheld devices; every model is automatically converted to it and validated against it before release.

MLflow logo

MLflow

Provides the versioned model registry, giving each guidance model one authoritative record from experiment through to fleet deployment.

Apache Airflow logo

Apache Airflow

Orchestrates the release pipeline end to end, from registry pull through conversion, on-device validation and the staged rollout gates.

Docker logo

Docker

Containerises the pipeline's build, conversion and validation stages so every release is produced in an identical, reproducible environment.

Kubernetes logo

Kubernetes

Runs the pipeline services and rollout backend, keeping the release machinery available while staged deployments progress across the fleet.

AWS IoT Greengrass

Carries models over the air to the handheld devices, managing staged deployment groups and executing rollbacks at the device edge.

Grafana logo

Grafana

Displays fleet version telemetry and rollout health, showing exactly which devices run which model as each release widens.

PostgreSQL logo

PostgreSQL

Stores the pipeline's audit trail and fleet state, the same records the automated release documentation for the quality team is generated from.

Ready to Build your Medical Devices Business with MLOps & Model Deployment

Ask Byte

Ask Byte

Typically replies instantly

just Now

Hi! I'm OrganByte's assistant. How can I help you today?

AI-generated content may be incorrect


OrganByte

Building innovative software solutions that transform businesses and drive digital success.

© 2026 YourCompany. All rights reserved.