
Fine-Tuned LLM Coaching Engine for a Wellness Subscription App
Let's Connect
Overview
What we built
A subscription wellness-coaching startup with roughly 40,000 members had hit a wall on how many clients each coach could handle. We fine-tuned a large language model to draft coaching replies in the programme's own voice, so coaches could support far more members without working longer hours.
In plain terms: each human coach could manage only about 200 clients before quality began to slip, and replies to member check-ins took hours to arrive. Coaching quality also varied noticeably from one coach to the next, so two members with the same question could get different guidance depending on who answered. Membership was growing fast enough that hiring at the same pace would have doubled payroll within a year, a cost the subscription business could not absorb.
We fine-tuned a large language model on thousands of anonymised, coach-approved conversation excerpts, so it drafts replies and habit nudges that sound like this programme's own coaches, not a generic assistant. Every draft passes through a coach in a unified inbox before it reaches a member, and sensitive topics route straight to a human with no AI draft attached. Average coach caseload grew from 200 to 470 members with no drop in satisfaction scores, and median reply time fell from 9 hours to under 30 minutes.
The Problem
Coaching couldn't scale
The coaching model relied entirely on human attention, and human attention does not scale the way a subscription membership does. With roughly 40,000 members on the books, each coach could take on only about 200 clients before the quality of their guidance began to suffer. Growth in membership therefore meant growth in headcount, on a curve the business could not sustain without hiring at a pace that would have doubled payroll within a year.
Speed suffered along with capacity. A member who wrote in with a question about a missed workout or a plateaued goal could wait hours for a reply, by which point the moment for encouragement had often passed. Coaches working through a backlog of check-ins had less time to think carefully about each one, so replies grew shorter and more generic under pressure.
Consistency was the quieter problem. Coaching quality varied noticeably between coaches, so a member's experience depended more on who happened to answer than on the programme itself. That variation was invisible in any single conversation but added up across 40,000 members into a service that did not feel the same for everyone paying for it.
Caseload ceiling
Each coach could manage only about 200 clients before quality began to slip, capping how many of the 40,000 members any coach could properly support.
Slow check-in replies
Replies to member check-ins took hours to arrive, so the encouragement or correction a member needed often landed well after the moment it mattered.
Inconsistent coaching quality
Coaching quality varied noticeably between coaches, meaning two members asking the same question could receive noticeably different guidance depending on who replied.
Unsustainable hiring curve
Hiring coaches at the pace membership was growing would have doubled payroll within a year, a cost the subscription business could not sustain.
What it was costing them
Every member stuck in a slow reply queue was a subscriber whose motivation was quietly draining away, and every inconsistent answer chipped at trust in the coaching itself. With the business needing to double payroll within a year just to keep pace with 40,000 members, growth and coaching quality were pulling in opposite directions, and neither could be fixed by hiring alone.
The Solution
Coach-in-the-loop fine-tuned LLM
We built the coaching engine around the programme's own voice rather than a generic chatbot personality. The model was fine-tuned on thousands of anonymised, coach-approved conversation excerpts, so its drafts already sound like this programme's coaching style rather than a stock assistant's. That grounding mattered because members trust their coach's voice, and a mismatch would have undermined the whole relationship.
Every draft still passes through a human. Coaches review, edit and send each message from a unified inbox, so the model proposes and a coach decides, never the other way round. Sensitive topics are automatically routed straight to a human with no AI draft attached at all, keeping the model out of conversations where judgement matters most.
We tuned the model in rounds rather than treating the first version as finished. Each round used more coach-approved excerpts and coach feedback on earlier drafts, so the assistant's replies needed less editing over time and coaches spent their attention on approval rather than rewriting from scratch.
Key decisions
Fine-tune on coach-approved excerpts
Training data came from thousands of anonymised, coach-approved conversation excerpts, so the model's drafts reflect this programme's own coaching voice rather than a generic style.
Coach review on every message
Coaches review, edit and send every message from a unified inbox, keeping a human decision between each AI draft and the member who receives it.
Route sensitive topics to humans
Sensitive topics are automatically identified and routed straight to a human, with no AI draft attached, keeping the model away from conversations that need direct human judgement.
Draft habit nudges, not just replies
The model drafts habit nudges as well as check-in replies, giving coaches a starting point for proactive outreach, not only reactive answers to questions.
Iterate across tuning rounds
The model was tuned across multiple rounds using coach feedback, so the share of drafts sent with minor or no edits kept rising round over round.
Measurable Impact
What changed after launch
Caseload and speed both moved. Average coach caseload grew from 200 to 470 members with no drop in satisfaction scores, meaning coaches supported well over double their previous client count without members feeling the service had thinned out. Median reply time to member check-ins fell from 9 hours to under 30 minutes, so encouragement now arrives while it still matters.
Trust in the drafts grew alongside their use. By the second tuning round, 88% of AI-drafted replies were sent with minor or no edits, showing coaches trusted the model's voice enough to use it as written rather than rewrite it. 30-day member retention improved by 14% across the two quarters after launch, evidence that faster, more consistent coaching kept members subscribed longer.
Coach caseload
About 200 members per coach
470 members per coach, same satisfaction
Reply time
Hours to reply to check-ins
Under 30 minutes, down from 9 hours
Draft quality
No AI drafting in the workflow
88% of drafts sent with minor or no edits
Member retention
Retention before the coaching engine launched
30-day retention up 14% post-launch
Headline results
Average coach caseload grew from 200 to 470 members with no drop in satisfaction scores
Median reply time to member check-ins fell from 9 hours to under 30 minutes
88% of AI-drafted replies sent with minor or no edits after the second tuning round
30-day member retention improved by 14% across the two quarters after launch
Tech & Tools Used
What powered the build
Every tool below earned its place in this engagement. Here is the part each one played.
Python
Runs the fine-tuning and inference pipeline that turns anonymised conversation excerpts into a model tuned specifically on this programme's coaching voice.
OpenAI Fine-Tuning API
Handles the fine-tuning runs on coach-approved excerpts, producing the successive tuning rounds that steadily reduced how much editing coaches needed to do.
Node.js (NestJS)
Serves the backend that manages coach inboxes, routes sensitive topics away from the model, and joins each AI draft to the right member conversation.
React Native
Powers the mobile coaching app members use to send check-ins and receive replies, keeping the coaching experience inside the app members already use daily.
PostgreSQL
Stores member profiles, conversation history and coach edits, giving the fine-tuning pipeline a reliable source of approved excerpts to train on.
pgvector
Indexes past conversation excerpts so similar member questions can be retrieved as context when the model drafts a new reply or habit nudge.
Redis
Caches active coach inbox state and in-flight drafts, keeping the review interface responsive as coaches move quickly between member conversations.
AWS ECS
Runs the coaching engine's backend services in production, scaling capacity as coach caseloads and member messaging volume grew.
Mixpanel
Tracks reply times, edit rates and retention across the member base, giving the team the data that showed each tuning round improving draft quality.
Ready to Build your Health & Wellness Business with LLM Integration & Fine-Tuning
Ask Byte
Ask Byte
Typically replies instantly
just Now
Hi! I'm OrganByte's assistant. How can I help you today?
AI-generated content may be incorrect

