Hero
Data & Knowledge AI
Synthetic & Augmented Data
Fashion Resale & E-Commerce

Synthetic and Augmented Training Data Pipeline for a Fashion Resale Platform's Listing AI


Let's Connect

Overview

What we built

A fashion resale platform processing thousands of secondhand garments daily across 11 hubs ran its listing AI on training data that could not keep pace, badly under-representing rare brands and damage types. We built a synthetic and augmented training-data pipeline that keeps the model fed without relying on manual labelling alone.

In plain terms: a fashion resale platform handling thousands of garments a day across 11 hubs had built an automated listing model to identify brand, category and condition, but the training data behind it could not keep up. Rare brands and damage types were badly under-represented, manual labelling costs kept climbing as the platform tried to close the gap by hand, and misclassified listings were driving returns and constant repricing work for a model that was only as good as its thinnest categories.

We built a synthetic and augmented training-data pipeline: standardised studio capture rigs feed an augmentation service generating lighting, background and occlusion variants, while a 3D rendering workflow composites synthetic examples of the rare garment categories and defects that real photos rarely captured. The labelled and synthetic training corpus grew 6x in 4 months at roughly one-third the previous per-image cost, attribute classification accuracy on long-tail brands and categories improved from 71% to 88%, manual listing corrections per 1,000 intake items fell 43% across all 11 hubs, and average time to publish an intake garment dropped from 9 minutes to under 4.

The Problem

Listing AI starved of data

Processing thousands of secondhand garments a day across 11 hubs put real pressure on the listing AI identifying each item's brand, category and condition, and the model was only as reliable as the data it had been trained on. That data was thin exactly where it mattered most: rare brands and damage types, the cases a lister most needed help with, were badly under-represented in the training set.

Closing that gap by hand did not scale. Manual labelling costs kept climbing as the platform tried to collect more examples of the categories it lacked, and the effort never quite caught up with how quickly new rare brands and damage patterns appeared in the intake stream. Every under-represented category stayed under-represented, no matter how much labelling budget went into it.

The consequences showed up downstream, in listings themselves. Misclassified items, wrong brand, wrong category, wrong condition, drove returns from buyers who received something other than what was listed, and drove constant repricing work as staff corrected listings the model had got wrong. The thin spots in the training data became visible, expensive problems on the live platform.

Thin rare-category data

Rare brands and damage types were badly under-represented in the training data, exactly the cases where listers most needed the model's help.

Climbing labelling costs

Manual labelling costs kept climbing as the platform tried to collect more examples by hand, without ever closing the gap in rare categories.

Misclassified listings

Misclassified brand, category or condition drove returns from buyers and constant repricing work for staff correcting the model's mistakes.

Scale outpacing labelling

Thousands of garments a day across 11 hubs meant new rare brands and damage types kept appearing faster than manual labelling could cover them.

What it was costing them

Every under-represented category meant a listing model guessing where it should have known, and every wrong guess turned into a return or a repricing task somebody had to handle by hand. Manual labelling costs kept climbing without closing the gap, so the platform was paying twice: once for labelling effort that never caught up, and again for the misclassified listings that effort failed to prevent across all 11 hubs.

The Solution

Synthetic training data pipeline

We built a synthetic and augmented training-data pipeline so the model no longer depended entirely on how many real examples a lister happened to photograph. Standardised studio capture rigs feed an augmentation service that generates lighting, background and occlusion variants from each real garment, multiplying useful training examples out of the images already being captured.

For the categories real photos rarely captured well, rare brands and specific defects, we added a 3D rendering workflow that composites synthetic examples directly. That gave the model coverage of long-tail cases without waiting for enough real secondhand garments of that exact type to pass through intake.

The pipeline stays current rather than being trained once and left. Versioned datasets flow into scheduled retraining, so the model keeps improving on a cadence, and lister corrections loop back automatically as fresh labelled samples, meaning every mistake caught on the live platform becomes training data instead of a one-off fix.

Key decisions

01

Standardise studio capture

Standardised studio capture rigs gave the augmentation service a consistent, high-quality base to generate lighting, background and occlusion variants from.

02

Augment from real images

An augmentation service multiplies each captured garment into lighting, background and occlusion variants, growing the usable training set without new photography.

03

Render synthetic rare cases

A 3D rendering workflow composites synthetic examples of rare garment categories and defects, covering long-tail cases real photos rarely captured.

04

Schedule retraining on versioned data

Versioned datasets flow into scheduled retraining, keeping the listing model current as the training corpus grows rather than training it once.

05

Loop lister corrections back in

Corrections listers make on misclassified items feed back automatically as fresh labelled samples, turning live mistakes into future training data.

Measurable Impact

What changed after launch

The training corpus stopped being the bottleneck. The labelled and synthetic corpus grew 6x in 4 months at roughly one-third the previous per-image cost, and attribute classification accuracy on long-tail brands and categories improved from 71% to 88%, closing the gap that used to sit exactly where listers needed the most help.

Those gains reached the live platform quickly. Manual listing corrections per 1,000 intake items fell 43% across all 11 processing hubs, and average time to publish an intake garment dropped from 9 minutes to under 4, as fewer listings needed a human to catch what the model had misclassified.

Training corpus

Growing slowly, thin on rare categories

Grew 6x in 4 months, synthetic-augmented

Classification accuracy

71% on long-tail brands and categories

88% on long-tail brands and categories

Listing corrections

High correction rate across 11 hubs

Down 43% per 1,000 intake items

Time to publish

9 minutes per intake garment

Under 4 minutes per intake garment

Headline results

Labelled and synthetic training corpus grew 6x in 4 months at roughly one-third the previous per-image cost

Attribute classification accuracy on long-tail brands and categories improved from 71% to 88%

Manual listing corrections per 1,000 intake items fell 43% across all 11 processing hubs

Average time to publish an intake garment cut from 9 minutes to under 4

Tech & Tools Used

What powered the build

Every tool below earned its place in this engagement. Here is the part each one played.

Python logo

Python

Ties the pipeline together, coordinating capture ingestion, augmentation, synthetic rendering and scheduled retraining into one automated flow.

PyTorch logo

PyTorch

Trains and retrains the listing AI's classification models on the growing labelled and synthetic corpus.

Albumentations

Powers the augmentation service, generating the lighting, background and occlusion variants from each standardised studio capture.

Blender logo

Blender

Runs the 3D rendering workflow that composites synthetic examples of rare garment categories and defects.

OpenCV logo

OpenCV

Handles image processing steps around capture and augmentation, preparing garment images for the augmentation service and the training pipeline.

Label Studio

Where lister corrections and manual labelling happen, feeding fresh labelled samples back into the versioned training datasets.

Apache Airflow logo

Apache Airflow

Schedules the retraining runs against versioned datasets, keeping the listing model current as new labelled and synthetic data arrives.

MLflow logo

MLflow

Tracks model versions and training runs, giving the team a record of how classification accuracy on long-tail categories improved over time.

PostgreSQL logo

PostgreSQL

Stores metadata for garments, labels and dataset versions, tying capture, corrections and retraining runs together.

AWS S3

Holds the studio captures, synthetic renders and versioned training datasets that feed the pipeline's scheduled retraining.

Ready to Build your Fashion Resale & E-Commerce Business with Synthetic & Augmented Data

Ask Byte

Ask Byte

Typically replies instantly

just Now

Hi! I'm OrganByte's assistant. How can I help you today?

AI-generated content may be incorrect


OrganByte

Building innovative software solutions that transform businesses and drive digital success.

© 2026 YourCompany. All rights reserved.