Customer Success: What Is Actually Happening
A dashboard flags 120 accounts as low engagement. Maya has six people and knows most of those flags are wrong, so she sits down and labels what is really going on — and those labels become the truth a model will learn from.
What you'll learn
- Explain what ground truth is, and why a model can only ever be as good as the labels it is given
- Separate three questions dashboards blur together: is this at risk, what is causing it, and can we act this quarter
- Recognise correlation dressed up as causation, and the cost of chasing it
It is a Tuesday, and Maya has a spreadsheet on one screen and a call list on the other. The spreadsheet came out of the customer success platform overnight and it says that 120 of Northwind’s 400 accounts have “low engagement”. Her team is six people. Six people cannot phone 120 companies, and even if they could, Maya already knows something the spreadsheet does not: a good number of those 120 are perfectly happy.
She knows it because she has spoken to them. The law firm on row 47 goes quiet every August and renews early anyway. The manufacturer on row 12 is using the product constantly and is furious. Until today that knowledge has lived in her head and in a hundred call notes. Now it has to come out in writing, in a form a machine can learn from — because Sofia has asked for a labelled sample, and whatever Maya writes down becomes the definition of “at risk” for the rest of this project.
A definition arrives; judgement is added; a set of verified examples moves on. Everything the model later believes about churn starts in the middle box.
What lands on Maya’s desk
What arrives is Sofia’s product requirements document from module 4, and with it constraints Maya cannot renegotiate. Sofia has already decided what an alert must contain — an account, a risk level, a top reason and a suggested next step — and has already chosen the four early-warning indicators the system will watch. Behind that sit Daniel’s mandate and Omar’s spending limit, which between them decided this project gets 26 weeks and not thirty-six.
The request itself is one line: label a sample of accounts so the data team has verified examples to learn from. It sounds administrative. It is the most consequential task in the first half of this project, because the PRD says what an alert should look like and nobody has yet said what a correct alert would be. Maya is being asked to supply the answer key.
What a Customer Success Lead actually does
If sales is the job of getting a signature, customer success is the job of making the signature worth repeating. Maya’s team runs onboarding, watches adoption, handles the awkward conversation when a sponsor leaves, and owns the renewal date. At Northwind, where the average contract is $24k a year and winning a new customer costs $25k, keeping an existing customer is worth about the same as winning a new one and is considerably easier.
The superpower she brings to an AI project is not analytical. It is context: she knows which quiet account is a crisis and which is a firm on holiday, because she has been on the phone to both. That knowledge is unstructured, undocumented and held entirely in the heads of six people. Every organisation has this asset. Almost none treat it as data.
The vocabulary of this desk
- Labelled data
- Examples where a human has recorded the true answer — “this account left for product reasons”, “this one was always fine”. Models learn from labels and nothing else.
- Ground truth
- The agreed record of what really happened, against which any prediction is judged. If it is wrong, every accuracy number computed later is measuring the wrong thing.
- Correlation vs causation
- Low usage goes together with churn. But court recess also causes low usage, with no churn at all. Confusing the two means chasing healthy customers while real leavers walk out.
- Actionable
- An alert someone can actually do something about this quarter. “Thornbury has low usage” is true and useless — they are locked into year one of three.
The software on Maya’s desk
The dashboard says which accounts look wrong. The other three screens say whether they are.
Maya’s morning moves from the general to the specific, and the software is arranged the same way. Power BI shows the shape of the problem across the whole customer base. Gainsight — a customer success platform — narrows that to a ranked list of accounts by health score, which is where the famous 120 flags come from. Both of those screens are quantitative, and both can be confidently wrong.
The correction lives in the last two. Salesforce, the CRM, holds one page per customer: who the contacts are, what they pay, when the contract ends. Zendesk holds what they have actually been saying — and it is here that Corvex Manufacturing stops being a healthy-looking heavy user and becomes a company with twenty-three open tickets about a feature that has been broken since March. This is the whole argument of module 5 rendered as software: the score is in one system and the reason is in another, and only a human currently opens both.
The software on this desk
- Power BI
- The reporting tool where the churn trend was first noticed — the picture that started the entire project.
- Gainsight
- A customer success platform: health scores, renewal dates, playbooks. It ranks accounts by engagement, which is useful and not the same as being right.
- Salesforce (the CRM)
- Customer relationship management: one page per customer holding contacts, contract value and renewal date. The system of record for who a customer is.
- Zendesk
- The support desk. Holds the tickets — the closest thing to a transcript of how the relationship is actually going.
The three decisions Maya makes
Deciding to label by hand at all
The fastest route to a training dataset is to pull “cancelled: yes/no” from the billing system. It is free, it covers all 400 accounts, and it is available this afternoon. Maya argues against it, and she is right.
What is being decided here is what the model will treat as reality. A model has no access to the world; it has access to your labels. It cannot visit Hale & Porter and notice the office is empty because the courts are in recess. It sees a number falling and a column marked true or false, and builds its entire understanding of churn from the relationship between them. Billing data records the outcome and destroys the reason — so it teaches nothing about the difference between a customer who is bored and a customer who is blocked.
Maya therefore decides to hand-label a small sample properly rather than auto-label everything badly. She takes six accounts from the flagged list, opens every evidence drawer — call notes, tickets, billing history, surveys, renewal history — and writes down what was really happening and why. The one-line summary is what a dashboard sees; the evidence underneath is what a human knows, and the point of the exercise is to move the second into the first.
Why a small careful sample beats a large careless one
Labels are the ceiling, not the floor. A model trained on 400 sloppy labels cannot outperform the sloppiness; a model tested against 60 careful ones will at least tell you the truth about itself. Accuracy against a bad answer key is not accuracy.Deciding what “at risk” actually means
This is the decision that changes the shape of the product. Maya realises, working through the accounts, that the dashboard is collapsing three separate questions into one number: is this account at risk, what is causing it, and could we do anything about it this quarter? Only when all three line up is an alert worth a phone call. Watch her apply that to the six.
Bramford Logistics is the account the whole project exists for. Their Ops Director — Maya’s champion, the person who bought the thing — left three weeks ago, and the replacement is “reviewing all software spend”. Weekly active users have fallen from 61 to 24 since June, the renewal is in nine weeks, and they have ignored the last two satisfaction surveys. Support is quiet: two minor tickets, both resolved. Risk: high. Cause: sponsor loss. Actionable: yes, urgently, and it needs an executive conversation this month rather than a friendly check-in. The quiet support queue is not reassurance; people who have stopped caring stop complaining.
Hale & Porter LLP looks nearly as alarming and is completely fine. Logins are down 45% in August — and were down in August last year, and the August before. A partner mentioned on a call that half the firm is on leave for court recess. They have renewed early twice without negotiating and scored 8/10 in June: “does what we need.” Risk: none. Cause: the calendar. This is where correlation and causation part company. Across 400 accounts, low usage genuinely does predict cancellation; the pattern is real. But the pattern is not the mechanism. Usage falls because people disengage, and usage falls because people are on holiday, and in a chart the two are identical. Ring Hale & Porter to save them and you spend a call, and some credibility, telling a happy customer you have noticed they are underperforming.
Corvex Manufacturing breaks the model most people carry in their heads, which is that risk means disengagement. Corvex use the product constantly — they run their production line on it — and they are furious. There are 23 open tickets, nearly all about the same export feature failing since the March release, and their last survey was 3/10: “fix it or we leave.” Risk: high. Cause: a product defect. Actionable: yes, but the right action is an engineering escalation with a date attached, and offering this account a discount would insult them. The distinction between risk and cause is not academic; it selects the response.
Juniper Health is the one nobody enjoys writing down. Their card has failed twice this quarter and both invoices were paid late. Of 50 licensed seats, 12 have logged in this year — they bought for a team that was then reorganised away. The survey says 6/10, “we’re just not using most of it”. Risk: high, renewal in three months. Cause: they are paying for capacity they do not need. The honest fix is a smaller contract, which means Maya is proposing to reduce revenue deliberately in order to keep a customer who would otherwise leave entirely.
Aldgate Media is genuinely wobbly and genuinely not urgent. Four months in, they cancelled two of three setup sessions and never switched on the workflow module they bought the product for. The sponsor is keen but cannot get his team to change how they work, and renewal is eight months away. Risk: moderate. Cause: onboarding never finished. Actionable: yes, as a nudge and a rescheduled training session, not a five-alarm call.
Thornbury County teaches the third question. Usage is low and flat, they use a fraction of the modules, and they are 22 months from renewal in year one of a three-year contract signed through formal procurement. Risk on paper: high. Actionable this quarter: no — public bodies here rarely exit mid-term. Putting Thornbury at the top of a ranked list every morning does not save them; it consumes the slot Bramford needed.
Deciding that a score without a reason is not an alert
The third decision is a rule, and it survives all the way to the end of the course: every alert must carry a cause. Not a probability, not a colour, not a health score out of a hundred — a stated reason, drawn from the same vocabulary Maya has just used. Sponsor loss. Product failure. Over-licensing. Failed onboarding. Seasonal. Contractually locked.
Maya insists on this because she manages the people who have to act. A rep handed “Corvex: 82% risk” opens a blank call with a customer who has told them, 23 times, exactly what is wrong. A rep handed “Corvex: high risk — product defect, export feature, since March” opens with an apology and a fix date. The score tells you who to ring. Only the cause tells you what to say, and only the cause makes it possible, later, to check whether the action matched the problem.
Where this goes wrong
Here is how it usually happens. The labelling exercise is scheduled, then squeezed, because customer success is busy doing customer success and this is “the data team’s project”. Somebody helpfully offers to save Maya the trouble by exporting cancellations from billing. Nobody objects, because objecting means arguing that six people’s phone calls are more authoritative than a system of record.
Five months later the model is live, and every August it lights up with law firms — confidently, because it was taught at scale that quiet accounts are dying accounts and has no way of knowing anyone ever thought otherwise. Meanwhile Corvex, heavy users and so never flagged, cancels without appearing on a single list.
The trap
A model does not have common sense to fall back on. It cannot notice that a rule it learned is absurd. Whatever judgement you compress into the labels is the judgement it will apply, uniformly, to every account, at three in the morning, forever.What Maya hands on
What leaves her desk is small and unglamorous: six accounts labelled with a risk level, a stated cause and an honest note on whether anything can be done this quarter, plus a written definition of actionable and the rule that no alert ships without a reason attached. It goes to Ben, the Data Analyst, in module 6, with the message: here is what churn looks like from the front line — now tell me where it has been hiding in the numbers.
The constraints travel a long way. In module 8, Priya’s model learns from and is judged against these labels, and can be no more correct than the answer key it is scored on — so Maya’s care this Tuesday sets the ceiling on Priya’s accuracy in December. In module 10, Kofi has to build the alert with the cause field mandatory rather than optional, a requirement that comes from here and not from engineering. And in module 11, Jade’s playbook branches on Maya’s causes, because “call them” is not a plan and “escalate the export bug, offer a smaller contract, reschedule the onboarding” is.
The bottom line
A model has no access to reality — it has access to your labels, and it will apply the judgement inside them at scale and without doubt. Maya’s contribution is to separate three things a dashboard blurs together: is this at risk, what is causing it, and can we act this quarter. The most important dataset in this project was created by a person with a phone and good judgement.Spot the label
Read each account, decide what you would write down, then tap the card to see what Maya wrote.
Quick check
1. Why does Maya refuse to let the data team label accounts from billing records alone?
2. Thornbury County has low usage and is genuinely under-adopting. Why does Maya label them "not actionable"?
3. What does Maya's rule that "every alert must carry a cause" change downstream?