Did It Work — And What Happens Now?
Nadia leaves a quarter of the at-risk accounts alone on purpose, refuses to call the result early, and hands Daniel a number small enough to be true — then the board decides what to scale, what to fix, and what to stop.
What you'll learn
- Design a comparison that separates a real effect from everything else that changed
- Pre-register one outcome metric and hold the end point when the good news arrives early
- Judge a whole system against the original mandate, and size the next bet to the evidence
It is a Monday, and Nadia has two tabs open. The first is the pilot dashboard, refreshed by somebody in the executive suite roughly every forty minutes. The second is an empty document titled Evaluation design, which she intends to finish and circulate before the model sends its first alert on Wednesday morning.
The order matters more than anything else she will do. In four weeks somebody senior is going to look at the first tab and want to announce something, and by then it will be far too late to decide what would have counted as success. Nadia’s entire contribution to this project is to write that down while nobody yet knows the answer — because retention at Northwind is about to improve, and at least four different things will be able to claim the credit.
Actions go in; a comparison comes out. This is the only desk in the chain whose job is to make the good news smaller and truer.
What lands on Nadia’s desk
From Jade in module 11: the logged actions. Which account, which alert, which response, on which day, with the reason — and, crucially, which accounts were deliberately left alone. They exist because Kofi spent his last three sprint points on one-click logging in module 10 rather than asking six busy people to fill in a form afterwards. That is the difference between six weeks of evidence and six weeks of anecdote.
From Omar in module 3: the bar. A minimum 12% reduction in churn among the at-risk accounts the programme contacts, inside a $250k envelope, both agreed in week two while agreeing was cheap. From Elena in module 9: the quiet enabler. Predictions are kept for twelve months, which is the only reason a score generated in one month can be laid beside what that customer actually did three months later. Delete nightly and this module does not exist.
And one complication, which is really why this module exists. The broken export feature — the one Corvex Manufacturing raised twenty-three tickets about — was fixed in week three of the pilot, paid for out of the $100k Daniel gave product engineering in module 2. Retention was going to improve a little whether or not the model ever worked. If nobody plans for that, the model is handed a win it did not earn, and the company scales a programme on the strength of somebody else’s bug fix.
What an experimentation lead actually does
A product analytics and experimentation lead designs measurement so that the result means something afterwards. Not builds dashboards — anybody can build a dashboard, and a dashboard will happily show you a rising line without ever telling you why it rose. The craft is in the comparison, and the whole discipline reduces to one question asked at the right time. Compared to what? Retention improved — compared to what? Everyone in a business can produce a number; almost nobody can produce one that survives that question. Nadia’s job is to arrange, in advance, for the answer to be a good one.
The vocabulary of this desk
- Counterfactual
- What would have happened anyway, without your programme. You can never observe it directly, which is why you have to build something that stands in for it.
- Holdout group
- At-risk accounts deliberately left alone, so you can see what happens without help. It feels cruel. It is the only way to know the programme works before spending millions scaling it.
- Pre-registration
- Writing down the one metric that decides success, and the date you will read it, before the data arrives.
- Statistical power
- The ability to detect a real effect if there is one. It comes from the number of outcomes you observe, not the number of accounts you enrol.
- Guardrail metric
- A second measure watched to check the win was not bought too expensively, or at somebody else’s cost.
- Peeking
- Checking early and stopping when it looks good. Randomness produces streaks; stopping on one enshrines luck as strategy.
The software on Nadia’s desk
Two groups, one comparison, and a chart that decides whether there is a next project.
Amplitude (or a similar behavioural analytics tool) tracks what the two groups actually did — the accounts that were worked and the holdout that deliberately was not — over the full quarter. Keeping those groups clean is more of the job than the statistics that follow; a rep quietly helping a holdout account out of kindness is the commonest way a good experiment dies.
The analysis itself is done in Python, because it needs things a dashboard cannot do: retention curves, confidence ranges, and an honest attempt to separate the effect of the programme from the effect of the export fix that shipped mid-pilot. Looker then carries the running pilot dashboard that leadership refreshes rather too often, which is exactly why the interim numbers on it are labelled as interim. And the result finally travels, like every other decision in this course, as a PowerPoint slide: before and after, with the uncertainty shown rather than hidden. It is the same tool Daniel used to ask for the money in module 2, which is a fitting way for the chain to close.
The software on this desk
- Amplitude
- Behavioural analytics. Tracks what treatment and holdout accounts actually did, which is the raw material of any honest comparison.
- Python
- Where the statistics happen — retention curves, significance, and separating the programme’s effect from everything else that changed.
- Looker / Power BI
- The pilot dashboard leadership watches. Useful, and the reason interim numbers must be labelled as interim.
- PowerPoint
- How the result reaches the board: before and after, with the uncertainty visible. The same medium the money was requested in.
The five decisions Nadia makes
The comparison group
Over the pilot, retention among at-risk accounts will improve, and Nadia can already name three explanations that will fit the data as well as hers. The export bug was fixed in week three, so Corvex and everyone like them stopped having a concrete reason to leave. The season turned; the annual August lull that makes Hale & Porter LLP look moribund every year passed, as it always does. And the reps got better, because six people doing the same difficult conversation sixteen times a week in October are simply better at it than they were in July. The programme is a fourth explanation, and without a comparison group it is exactly as well supported as the other three, and no more.
So Nadia holds back a random 25% of the accounts that cross the alert threshold. They are scored, they are ranked, they appear on nobody’s list, and nothing is done about them. Everything else about them is identical: same model, same weeks, same fixed export, same weather.
The word doing the work is random, not large. This is the part people get wrong. The tempting alternative — compare the accounts the reps contacted with the accounts they did not — costs nothing and looks like a comparison. It is worthless, because the reps chose: they rang the accounts they thought were saveable and quietly let go of the ones they had privately written off. Compare those two groups and you will have measured the reps’ optimism with great precision and your programme not at all. A random split is the one procedure that makes two groups alike in every respect, including the ones nobody thought to record.
The other tempting alternative is comparing this quarter with last quarter — a comparison to a different world, one with a broken export feature in it, at a different point in the year. It bakes the product team’s fix into the model’s result. Flattering, and false.
Say this part out loud
A holdout means deliberately not helping customers who might have been helped. It is uncomfortable and it should be — Nadia says so to Maya and Jade rather than burying it in a methodology appendix. The defence is proportion: twenty-two accounts left alone for one quarter is the price of knowing whether it is right to do this for all four hundred, for years, at several hundred thousand dollars. Skipping the holdout does not spare anyone. It moves the harm somewhere nobody can see it.How long it runs
Interventions run for six weeks. Nobody is allowed to conclude anything for twelve — a full quarter — because churn is measured in renewal decisions, and renewal decisions arrive long after the phone call that was meant to influence them.
The pressure is to run three weeks and report. Three weeks would enrol plenty of customers, which sounds like a decent sample and is useless anyway, for the most misunderstood reason in business measurement. Statistical power does not come from how many accounts you enrol. It comes from how many outcome events you observe. Enrol two hundred accounts into a three-week window and you might see four renewals — and four events cannot distinguish a real effect from noise. The analysis will not refuse to produce a number; it will produce one, with an uncertainty range wide enough to include both “this saved the company” and “this made things worse”.
Even across a full quarter, about 88 distinct accounts cross Priya’s threshold and only about 20 renewal decisions land inside the window. Nadia says that at the start rather than discovering it at the end. It is why the final range will be wide, and why she will not pretend otherwise. The trade runs both ways: six months would be statistically lovely and organisationally fatal, because leadership will not hold a pilot open that long without a reading, and a perfect test nobody waits for teaches nothing either. A quarter is where the evidence and the patience meet.
One metric, named in advance
The primary metric is ninety-day retention of at-risk accounts, treatment against holdout. One number, written down, circulated and dated before the first alert goes out.
Why one? Because with several candidate metrics and no commitment, somebody will find the one that moved and present it. Not dishonestly — that is genuinely what searching a table feels like from the inside. There is always a segment, a week or a cut where the line goes up, and the person who finds it will believe they have discovered something rather than manufactured it. Pre-registration works precisely because it is done in ignorance: you decide what would count as success while you still do not know the answer. Afterwards, everyone is a genius at choosing metrics.
Why this number? Because it is the promise the business case made. Omar did not fund a programme to make phone calls. He funded a programme to keep customers, and priced it at 12%. The metric that decides a project should be the sentence the project was sold on.
Notice what gets demoted. Number of interventions completed is an activity metric — it counts effort, not outcome, and can be maximised by doing more work to no effect whatsoever. Ten thousand immaculate, cause-matched calls that save nobody would score magnificently on it. Model AUC is worse: 0.84 is a fact about the model, not about the business, and a perfect model whose alerts nobody acts on saves precisely no one. Both stay in the readout as context. Neither gets a vote.
The guardrails
Two secondary measures, chosen for one property: the pilot could plausibly move them.
The first is cost per save against the business case. A primary metric can nearly always be improved by spending more than the result is worth, and nothing else in the design would notice. Omar’s arithmetic put a saved customer at roughly $18k of gross profit a year; if the programme is spending $20k to save one, it is working beautifully and destroying value.
The second is satisfaction among contacted customers. Precision at Priya’s threshold is about two in five, so ten of every sixteen calls go to a customer who was fine — an accepted cost, but one that lands on people rather than on a budget line, and could quietly turn into irritation among healthy accounts nobody was worried about.
The rejected guardrail is the instructive one. Somebody always proposes total company revenue, on the grounds that it is what ultimately matters. It is — and it is useless here, because it moves for a hundred reasons unconnected to a pilot touching sixty-odd accounts: one large renewal, a pricing change, an enterprise deal slipping a fortnight. A guardrail that cannot be moved by the thing you are doing is not protecting anything; it is decoration with a chart attached.
Holding the end point when the news is good
In week four the email arrives, from the CEO’s chief of staff, and it is entirely reasonable in tone. Results look amazing — can we announce and roll out to all accounts next month?
They do look amazing. The treated group is running at what looks like a 41% reduction in churn against the holdout. Nadia knows three things about that number. It rests on a handful of renewal decisions, and early streaks are simply what randomness looks like from close up — flip enough coins and a run of heads always turns up somewhere. Stopping the moment results look good is not merely risky; it is wrong in a systematic direction, because you only ever stop early on a high, so peeking projects report effects inflated by roughly the amount that made stopping tempting. And that number does not stay a number: it becomes the line in the deck, then the assumption in next year’s plan, then the target somebody who was not in this room is measured against and fails to hit.
So Nadia does two things. She shares the interim figures openly, clearly labelled interim, with the range attached and a plain sentence saying the final number will very probably be smaller. And she holds the pre-registered end point. It is mildly awkward for about a week.
Why the rules were written down first
The entire purpose of deciding the rules in advance is that it will be inconvenient later. A design that only holds when the news is bad is not a design; it is a formality that has never been tested. Being the person who says “wait” while everyone else is drafting the announcement is not obstruction. It is the job.What the numbers actually said
At ninety days, churn among the treated at-risk accounts was 21%. Among the holdout, 26%. That is a 19% relative reduction in churn among contacted accounts, with a confidence range running from roughly 7% to 31%.
Nadia makes the board read that properly. The central estimate clears Omar’s 12% floor; the bottom of the range does not. With twenty renewal decisions, that is the honest width of what a quarter can tell you, and no further analysis narrows it — only more events would. Cost per save came in at about $6.2k against a customer worth $18k a year in gross profit, comfortably inside the case. Satisfaction among contacted customers was flat, including among the healthy accounts contacted in error — a guardrail doing its job by staying still.
Then the sentence that took her longest to write, and which is the bravest thing in this entire course:
Raw retention among at-risk accounts improved by far more than 19%. Most of that improvement also appeared in the holdout, which received no intervention at all — so most of it was the export fix and the turn of the season, not us.
That line is what the holdout bought. Without it Northwind would have reported something near a 34% improvement, credited it to the model, and scaled a programme built partly on somebody else’s bug fix. A good design does not only measure the effect; it gives the credit to the right people — here, the product engineers Daniel funded in module 2, whose fix has now appeared twice: once as a genuine improvement in retention, and once as the confound that very nearly stole the credit for it.
Where this goes wrong
The common failure has three stages and takes about a year. No holdout, because leaving customers unhelped felt wrong and the design got softened in a meeting where nobody wanted to be the person arguing for it. Then an early stop on a lucky streak, announced with a percentage that is real in the sense of having been calculated. Then rollout, at which point the effect reverts to its true size — because it was always its true size — and the promised savings never arrive.
What the organisation concludes is not “we measured badly”. It concludes that the model stopped working, or that the vendor oversold it, or, most damagingly, that AI does not work here. The next project inherits a leadership team that has been burned once and will not fund a holdout, which guarantees the same ending again.
What Nadia hands on
An experiment readout: the pre-registered design, treatment against holdout, the effect with its uncertainty stated rather than buried, the guardrail outcomes, cost per save against Omar’s case, and an explicit paragraph on what the programme cannot claim credit for. It goes to Daniel and the board with one sentence attached — here is what actually happened, with the uncertainty on it; it is enough to decide with.
The constraint she passes on is not a number but a boundary on what may be claimed. Nobody in that boardroom can now say the model saved the company, because the holdout ruled it out in writing, dated before anybody knew the answer.
Back in the same room, six months later
His dashboard sets actuals against the original mandate, line by line, and he walks it without decoration. The problem was a 15% revenue decline, $11.3m to $9.6m, with churn doubled from 9% to 18% on 400 accounts. All 400 are now scored nightly. Over the pilot, 88 crossed the risk threshold; 66 were worked and 22 deliberately were not. The measured churn reduction among contacted accounts is 19%, in a range of 7% to 31% — about four accounts retained that the holdout says would otherwise have gone, some $96k of recurring revenue and $72k of gross profit, on a rate that annualises to a little under $290k if it holds. Spend is $238k against the $250k envelope. Elapsed time is 28 weeks against the 26 everyone agreed to, and Daniel says the two weeks out loud rather than rounding them away.
Then the open risks, because a readout without them is a sales pitch. The model will drift, and Elena’s control means a named person owns noticing rather than a team owning it in the abstract. Eight leavers a week never reach any list at all, which is what a 43% recall looks like written in customers rather than percentages. The lower end of the effect range sits below the bar Finance set. And the economics are proven at sixty-six accounts, not at four hundred.
Sizing the next bet to the evidence
That sentence names the rarest executive skill in this course: proportionality. Scaling weak evidence and refusing to kill a failure are not opposite errors. They are the same error — acting out of proportion to what you actually know.
Three rows decide it. Did the measured result clear the bar Finance set? Yes, centrally, though not at the bottom of the range. Is the evidence strong enough to believe? Yes, because there was a random holdout and a pre-registered metric, which is why a sceptic can be answered rather than out-argued. Do the economics survive several times the volume? This is the one still open, because $6.2k per save was achieved by six trained people working sixteen alerts a week, and nobody has yet shown what that figure does at sixty.
So the board funds phase two in proportion. Roll out beyond the pilot team. Make the pipeline and the labelled data permanent, funded as infrastructure rather than as a project — a budget-line distinction that decides whether they still exist in two years. Keep measuring, with a smaller 10% holdout running indefinitely, because the alternative is never knowing again. And revisit the threshold as capacity grows, since sixteen alerts a week was always a staffing decision wearing a technical costume.
The ending nobody plans and most projects get
The commonest fate of a pilot is neither scaling nor stopping. It simply continues at pilot size forever — funded, staffed, mildly successful, deciding nothing, consuming a budget line whose owner has long since stopped asking what it is for. Nobody ever chooses that. It is what happens when nobody chooses.Judging the system, not the model
The last thing Daniel says is what he most wants the board to remember, and it is not about the model. The model’s AUC is 0.84: respectable, unremarkable, and nearly beside the point. What produced the 19% was a chain. Maya’s rule that every alert carries a cause, so a rep opened a conversation about unused seats rather than reading out a risk score. Priya’s threshold set from six people’s real capacity rather than from a curve. Kofi’s one-click logging, without which none of this could have been measured. Jade’s cause-matched playbook, which sent something different to Bramford Logistics than to Juniper Health. Swap any one of those for its lazy alternative and the same model produces nothing.
That is the test to apply to any project of this kind. A good model with a broken intervention process is a failed project — and it will be reported as a model failure, because the model is the visible part. Had Jade’s team fired the same panicked discount at all sixteen alerts, the readout would have shown a flat line, the board would have concluded the predictions were wrong, and the actual fault would have sat three desks downstream of the thing that got blamed.
What the chain adds up to
Look back at module 1 and the image this course opened with: a rope, not a relay. In a relay, once you have passed the baton your race is over. Here, every decision kept acting on people months after the person who made it had moved on to something else.
You can see Daniel’s budget split still working inside the final number, and the thread is worth tracing once. The $50k he protected for training decided how many accounts six people could genuinely work in a week. That capacity set Priya’s threshold at sixteen alerts. Sixteen alerts a week produced eighty-eight enrolled accounts and about twenty renewal decisions. Twenty events is what made Nadia’s range run from 7% to 31% rather than something tighter — which is what obliged the board to fund phase two carefully rather than confidently. A training budget agreed in week two, by a man thinking about headcount, set the statistical precision of the answer in week twenty-eight. Nobody designed that. It is simply what happens when decisions travel.
The same is true of the $50k for governance, which produced Elena’s twelve-month retention rule and therefore the possibility of measuring anything at all; and of the $100k for product engineering, which fixed the export bug and thereby both improved retention and very nearly took the credit for the model.
The bottom line
A working AI solution is not built by a data scientist. It is the result of coordinated decisions across strategy, finance, product, customers, analytics, engineering, science, governance, operations and measurement — a budget split, an economic floor, a scope line, a set of honest labels, four queries, a pipeline, a threshold, five controls, a playbook, six weeks of logged actions, one comparison group, and a board call made in proportion to the evidence. The model was one link in twelve. The chain is the project, and every link is somebody’s ordinary Tuesday.Spot the flaw
Read each claim, decide what is wrong with it, then tap the card.
Quick check
1. Why does Nadia insist the holdout be chosen at random rather than simply be large?
2. What does the holdout reveal about the export bug fix?
3. What does Daniel mean by judging the system rather than the model?
Certificate of Completion
This certifies that
Your Name
has successfully completed
An AI Project, End to End
Foundry Rock Learning
Credential ID
Verify at
🎓 Claim your verifiable certificate
Create a free account (or log in) to get a credential ID you can verify online and add to your LinkedIn profile in one click.
Create free account Log in🎓 Claim your verifiable certificate