Hero
Generative & Agentic AI
LLM Integration & Fine-Tuning
Health & Wellness

Fine-Tuned LLM Coaching Engine for a Wellness Subscription App


Let's Connect

Overview

What we built

A subscription wellness-coaching startup with roughly 40,000 members had hit a wall on how many clients each coach could handle. We fine-tuned a large language model to draft coaching replies in the programme's own voice, so coaches could support far more members without working longer hours.

In plain terms: each human coach could manage only about 200 clients before quality began to slip, and replies to member check-ins took hours to arrive. Coaching quality also varied noticeably from one coach to the next, so two members with the same question could get different guidance depending on who answered. Membership was growing fast enough that hiring at the same pace would have doubled payroll within a year, a cost the subscription business could not absorb.

We fine-tuned a large language model on thousands of anonymised, coach-approved conversation excerpts, so it drafts replies and habit nudges that sound like this programme's own coaches, not a generic assistant. Every draft passes through a coach in a unified inbox before it reaches a member, and sensitive topics route straight to a human with no AI draft attached. Average coach caseload grew from 200 to 470 members with no drop in satisfaction scores, and median reply time fell from 9 hours to under 30 minutes.

The Problem

Coaching couldn't scale

The coaching model relied entirely on human attention, and human attention does not scale the way a subscription membership does. With roughly 40,000 members on the books, each coach could take on only about 200 clients before the quality of their guidance began to suffer. Growth in membership therefore meant growth in headcount, on a curve the business could not sustain without hiring at a pace that would have doubled payroll within a year.

Speed suffered along with capacity. A member who wrote in with a question about a missed workout or a plateaued goal could wait hours for a reply, by which point the moment for encouragement had often passed. Coaches working through a backlog of check-ins had less time to think carefully about each one, so replies grew shorter and more generic under pressure.

Consistency was the quieter problem. Coaching quality varied noticeably between coaches, so a member's experience depended more on who happened to answer than on the programme itself. That variation was invisible in any single conversation but added up across 40,000 members into a service that did not feel the same for everyone paying for it.

Caseload ceiling

Each coach could manage only about 200 clients before quality began to slip, capping how many of the 40,000 members any coach could properly support.

Slow check-in replies

Replies to member check-ins took hours to arrive, so the encouragement or correction a member needed often landed well after the moment it mattered.

Inconsistent coaching quality

Coaching quality varied noticeably between coaches, meaning two members asking the same question could receive noticeably different guidance depending on who replied.

Unsustainable hiring curve

Hiring coaches at the pace membership was growing would have doubled payroll within a year, a cost the subscription business could not sustain.

What it was costing them

Every member stuck in a slow reply queue was a subscriber whose motivation was quietly draining away, and every inconsistent answer chipped at trust in the coaching itself. With the business needing to double payroll within a year just to keep pace with 40,000 members, growth and coaching quality were pulling in opposite directions, and neither could be fixed by hiring alone.

The Solution

Coach-in-the-loop fine-tuned LLM

We built the coaching engine around the programme's own voice rather than a generic chatbot personality. The model was fine-tuned on thousands of anonymised, coach-approved conversation excerpts, so its drafts already sound like this programme's coaching style rather than a stock assistant's. That grounding mattered because members trust their coach's voice, and a mismatch would have undermined the whole relationship.

Every draft still passes through a human. Coaches review, edit and send each message from a unified inbox, so the model proposes and a coach decides, never the other way round. Sensitive topics are automatically routed straight to a human with no AI draft attached at all, keeping the model out of conversations where judgement matters most.

We tuned the model in rounds rather than treating the first version as finished. Each round used more coach-approved excerpts and coach feedback on earlier drafts, so the assistant's replies needed less editing over time and coaches spent their attention on approval rather than rewriting from scratch.

Key decisions

01

Fine-tune on coach-approved excerpts

Training data came from thousands of anonymised, coach-approved conversation excerpts, so the model's drafts reflect this programme's own coaching voice rather than a generic style.

02

Coach review on every message

Coaches review, edit and send every message from a unified inbox, keeping a human decision between each AI draft and the member who receives it.

03

Route sensitive topics to humans

Sensitive topics are automatically identified and routed straight to a human, with no AI draft attached, keeping the model away from conversations that need direct human judgement.

04

Draft habit nudges, not just replies

The model drafts habit nudges as well as check-in replies, giving coaches a starting point for proactive outreach, not only reactive answers to questions.

05

Iterate across tuning rounds

The model was tuned across multiple rounds using coach feedback, so the share of drafts sent with minor or no edits kept rising round over round.

Measurable Impact

What changed after launch

Caseload and speed both moved. Average coach caseload grew from 200 to 470 members with no drop in satisfaction scores, meaning coaches supported well over double their previous client count without members feeling the service had thinned out. Median reply time to member check-ins fell from 9 hours to under 30 minutes, so encouragement now arrives while it still matters.

Trust in the drafts grew alongside their use. By the second tuning round, 88% of AI-drafted replies were sent with minor or no edits, showing coaches trusted the model's voice enough to use it as written rather than rewrite it. 30-day member retention improved by 14% across the two quarters after launch, evidence that faster, more consistent coaching kept members subscribed longer.

Coach caseload

About 200 members per coach

470 members per coach, same satisfaction

Reply time

Hours to reply to check-ins

Under 30 minutes, down from 9 hours

Draft quality

No AI drafting in the workflow

88% of drafts sent with minor or no edits

Member retention

Retention before the coaching engine launched

30-day retention up 14% post-launch

Headline results

Average coach caseload grew from 200 to 470 members with no drop in satisfaction scores

Median reply time to member check-ins fell from 9 hours to under 30 minutes

88% of AI-drafted replies sent with minor or no edits after the second tuning round

30-day member retention improved by 14% across the two quarters after launch

Tech & Tools Used

What powered the build

Every tool below earned its place in this engagement. Here is the part each one played.

Python logo

Python

Runs the fine-tuning and inference pipeline that turns anonymised conversation excerpts into a model tuned specifically on this programme's coaching voice.

OpenAI Fine-Tuning API logo

OpenAI Fine-Tuning API

Handles the fine-tuning runs on coach-approved excerpts, producing the successive tuning rounds that steadily reduced how much editing coaches needed to do.

Node.js (NestJS) logo

Node.js (NestJS)

Serves the backend that manages coach inboxes, routes sensitive topics away from the model, and joins each AI draft to the right member conversation.

React Native logo

React Native

Powers the mobile coaching app members use to send check-ins and receive replies, keeping the coaching experience inside the app members already use daily.

PostgreSQL logo

PostgreSQL

Stores member profiles, conversation history and coach edits, giving the fine-tuning pipeline a reliable source of approved excerpts to train on.

pgvector logo

pgvector

Indexes past conversation excerpts so similar member questions can be retrieved as context when the model drafts a new reply or habit nudge.

Redis logo

Redis

Caches active coach inbox state and in-flight drafts, keeping the review interface responsive as coaches move quickly between member conversations.

AWS ECS

Runs the coaching engine's backend services in production, scaling capacity as coach caseloads and member messaging volume grew.

Mixpanel logo

Mixpanel

Tracks reply times, edit rates and retention across the member base, giving the team the data that showed each tuning round improving draft quality.

Ready to Build your Health & Wellness Business with LLM Integration & Fine-Tuning

Ask Byte

Ask Byte

Typically replies instantly

just Now

Hi! I'm OrganByte's assistant. How can I help you today?

AI-generated content may be incorrect


OrganByte

Building innovative software solutions that transform businesses and drive digital success.

© 2026 YourCompany. All rights reserved.