Governance: The Conditions Under Which Yes Stays Safe
Priya's model works, and Sales has already asked for it. Elena spends one week deciding what the score is allowed to do, who may see it, how long it lives and who owns it when it drifts — and hands Kofi five conditions rather than a signature.
What you'll learn
- Separate what a model can technically do from what it is permitted to do
- Choose controls that change behaviour rather than controls that merely sound strict
- Write a permitted-use rule tight enough to refuse a powerful colleague
It is a Thursday, and Elena has Priya’s model card open on one screen and a blank assessment template on the other. The model card is good — honest about what the model gets right, honest about where it fails, clear that every score arrives with a reason attached. The planned use is written down: daily churn scores landing in the customer success platform, humans deciding what to do about them.
Underneath that, in the covering note, is one line that will take up most of her week. Sales have asked whether they can have access too, to prioritise upsells. Nobody is doing anything wrong. Sales have seen a useful thing and asked politely for it, which is exactly how a retention model becomes pricing infrastructure — not through a decision anyone would defend, but through a sequence of small reasonable ones nobody wrote down.
A capability arrives; a boundary is added; an approval with conditions moves on. Nothing Elena writes makes the model better — all of it decides whether the model stays deployable.
What lands on Elena’s desk
Three things come across from Priya in module 8. The model card, which sets out what the model looks at, how accurate it is on Maya’s labelled examples, and where it is known to fail. The intended use: daily scores, delivered into the tool customer success already work in, with a stated reason beside each account. And the note about Sales.
A good deal is already fixed and cannot be reopened here. Sofia refused automated retention emails in module 4 on the record, so the design already assumes a person makes the call. Maya’s rule from module 5 — every alert carries a cause — means the score is never a bare number, which is the single thing that makes meaningful human review possible rather than theatrical. And Daniel funded governance at $50k back in module 2, protecting it early rather than trimming it last. That is why this is a one-week review sitting inside the project rather than a three-week scramble in an external counsel’s queue, discovered in week twenty-one. Cheap insurance is only cheap if you buy it before the claim.
What a governance and risk lead actually does
Governance turns technically possible into appropriately used. In practice that means answering four questions nobody upstream had time for: what decisions may this score influence, who may see it, how long does the data live, and when must a human review before anything happens to a customer. It is judged strangely, because good governance is invisible — it reads as “the project shipped and nothing blew up”, which from the outside is indistinguishable from nothing having been at risk.
The craft is in choosing controls that change what actually happens rather than controls that sound strict. Blanket bans and rubber stamps are both theatre; one destroys the value, the other permits everything while producing paperwork. Elena is looking for the smallest set of rules that would genuinely stop the bad version of this project.
The vocabulary of this desk
- Permitted use
- The written boundary saying what the model is for. This score exists for retention outreach; using it to set prices or screen support queues is a different project needing its own review.
- Human in the loop
- A person reviews before action is taken. The model may put Bramford at the top of a list; only a human decides what happens to Bramford.
- Model drift
- The world changes and the model does not. A model trained before the March price rise slowly becomes wrong while still returning confident numbers.
- Least privilege
- Each role gets the narrowest view that lets it do its job. Access is a control, not a courtesy, and breadth is its cost.
- Data minimisation
- Use the least data that does the job — usage trends and payment flags, not everything the company happens to know about a customer.
The software on Elena’s desk
Access is a list, not a vibe — and the list is what the auditor reads.
OneTrust and tools like it exist to convert an uneasy feeling into a structured assessment: what data is used, for what purpose, on what basis, with what safeguard. That structure is what allows a review to be finished rather than endlessly reopened, and it is why Daniel’s $50k of governance funding buys a one-week formality instead of a three-week scramble.
Microsoft Purview scans systems to find and tag personal data, on the sound principle that you cannot protect data you do not know you hold. Entra ID — the identity system, formerly Azure Active Directory — is where role-based access stops being a policy and becomes something technically enforced: customer success sees the accounts they own, leadership sees aggregates, and the audit trail proves it. ServiceNow GRC holds the register tying each risk to its control, its named owner and the evidence it is genuinely operating. Notice that every one of these tools produces a record. Governance runs on the assumption that in two years somebody will ask what you decided and why, and only the written answer will count.
The software on this desk
- OneTrust
- Privacy assessment software. Turns ‘should we?’ into a documented set of named questions, answers and evidence.
- Microsoft Purview
- Discovers and classifies sensitive data across a company’s systems, so protection can be applied to data you know you hold.
- Entra ID (Azure AD)
- The identity system. Where role-based access is actually enforced, together with the audit trail proving who could see what.
- ServiceNow GRC
- Governance, risk and compliance: the register linking each risk to a control, an owner and the evidence it is working.
The six decisions Elena makes
What the score is allowed to decide
The Sales request is the first and most consequential. They want the churn score to help decide who gets discount offers — which means, stated plainly, quietly different prices for different customers based on a model’s guess about their loyalty.
What is being decided is whether capability implies permission. Models spread by reuse, not by launch. A score built for retention becomes a pricing input, then a support-triage input, then a factor in who gets a renewal call at all, and nobody ever made that decision — they each made a small one. Score-driven differential pricing also carries legal weight the model was never audited for, because Priya tested it for accuracy against churn, not for fairness in setting prices.
Elena’s control is a permitted-use rule binding the model to the purpose that was actually reviewed, with new purposes requiring a new assessment. Note what she does not do: she does not refuse Sales, and she does not delete the model to be safe. Governance by prohibition kills the value along with the risk. She gives the new use a door to knock on rather than a wall to climb — a review request is a healthy sign; a silent expansion is the failure.
How much authority the score has over a phone call
A rep might trust the score blindly and ring a healthy customer to “save” them, or might ignore context the model cannot possibly see — that the customer is mid-merger, that last month’s call went badly, that their champion has already privately promised to renew.
Elena’s control is human review: the score and its reason are advisory, the rep decides, and the rep records why. The tempting alternative is to hide the score entirely and show only call or don’t call, which sounds like it removes the risk of over-trust. It does the opposite. Hiding the reasoning does not create oversight; it creates obedience with extra steps. People trust — and correct — what they can see. A rep shown “high risk: sponsor loss, logins down 60% since June” can say no, I spoke to them Tuesday; a rep shown a green tick can only comply.
The line that governs every AI product
Automate the ranking, never the relationship. The model is good at deciding who to look at first. It knows nothing about what to say, and it will never know what happened on the last call.Who notices when the model starts getting worse
Nothing announces drift. The March price rise changes how customers behave; the export bug Corvex complained about twenty-three times gets fixed and stops predicting anything; a competitor launches. The model keeps returning confident numbers while quietly getting worse, and because the numbers still look like numbers, people keep acting on them for months.
The control is monthly performance monitoring, a threshold that triggers a retrain, and — the part that actually matters — a named owner. Not “the data team”. A person. The alternative that sounds most modern, retraining automatically every night, is wrong twice over: it absorbs whatever new bias has appeared in the data with nobody looking, and it makes the reasons shown to reps unstable, so the explanation an account carries on Tuesday contradicts Monday’s for no visible cause. Drift is inevitable; unowned drift is the problem, and unowned decay is the default outcome.
Who can see an individual customer’s score
The build could easily expose every account’s churn score to everyone with a login, and someone will argue for that on the grounds of transparency.
A churn score is a judgement about a relationship, made by a machine, about a company that never agreed to be judged. Spread widely, it leaks into conversations it should never steer: a pricing call where someone remembers the number, a support queue where a “disloyal” account waits longer, an opinion formed in a corridor. None of that is auditable. Elena’s control is role-based access — customer success sees operational detail for their own accounts, leadership sees aggregates — enforced in the platform rather than promised in a policy, because a policy nobody technically implements is a wish. Restricting it to the data team alone would be the mirror-image error: then the people who must act on the alerts cannot see them, and the product has no users.
How long a prediction should live
Twelve months, then aggregate. This decision sits between two failures rather than at one extreme, which is why it is the one most often made by accident.
Keep everything forever, because storage is cheap, and Northwind has built a permanent loyalty dossier on 400 customers — thousands of stale judgements generated by models that no longer exist, misleading to anyone who reads them and entirely discoverable if a dispute ever gets formal. Storage is cheap; liability is not. Delete nightly, and the opposite failure: Nadia cannot compare predictions to outcomes in module 12, so nobody can ever prove whether the thing worked, or whether it was fair to any particular kind of customer. Evaluation needs memory. Elena ties the period to a stated purpose so it can be defended: twelve months, to measure outcomes. That is a policy. “Indefinitely, storage is cheap” is an accident waiting for a subpoena.
The sentence itself
The last decision is a single line of text, and it is the one Elena spends longest on, because it is the sentence that will be quoted back at every future clever idea for this model. She writes:
Churn scores may inform retention outreach by customer success, with human review. Any other use requires a new assessment.
Look at what is packed into thirty words. It names the users — customer success, not “the business”. It names the purpose — retention outreach, not “customer management”. It names the safeguard — human review. And it names the route to change it, so expansion has a legitimate path instead of a quiet one.
Now compare the version that gets written when nobody is paying attention: scores may be used for any commercial purpose approved by a manager. It reads as governance. It is longer than Elena’s sentence and contains the word “approved”. But “approved by a manager” is not a boundary — it is a signature waiting to be collected, and in any organisation there is always a manager who will sign. Within a quarter the score is setting prices and screening support tickets, and every step of that was authorised.
The test Elena applies is simple and worth stealing. Could this sentence be used, by a junior person, to refuse a powerful colleague, without a meeting? If not, it is decoration. Governance survives in the specific: precision is what lets someone with no authority hold a line drawn by someone with authority.
Where this goes wrong
The common version in real companies is not a scandal. It is a calendar problem. The review is scheduled for after launch, because governance is thought of as a sign-off rather than a design input, and the project is busy. Then it happens, and discovers the automated email feature someone helpfully added in a sprint, and that scores are visible company-wide because that was the default, and that nobody owns monitoring — and the fixes are now retrofits into a live system with users and habits attached.
That project loses four months and, more expensively, its credibility. The organisation concludes that governance slows things down, when what slowed things down was doing it in the wrong order.
The trap
Approval and conditions are not the same deliverable. “Approved” arrives once and expires never. Conditional approval means the model is permitted while the conditions hold — which makes the approval revocable, and gives everyone downstream a reason to keep the controls working after launch, when the attention has moved on.What Elena hands on
What leaves her desk is a risk assessment, five controls and a permitted-use statement, entered in the register where anyone can find them, and approval to deploy — conditional on human review, access control and monitoring being built, not promised. It goes to Kofi, the ML Engineer, in module 10, with a sentence that decides how he plans his sprint: approved, with conditions, and the conditions are requirements now, not suggestions. Build them in; do not bolt them on.
The constraints travel far. Human review and role-based access take capacity in Kofi’s backlog before any feature does, which is why his release is smaller than Sofia hoped — governance is not free, it is prepaid. The twelve-month retention is what makes Nadia’s measurement possible in module 12; without it the whole project ends in opinion. And the permitted-use line is the sentence read back to Sales within a month, when the discount idea returns. That is not the rule failing. That is the rule working, precisely as designed, without a meeting.
The bottom line
Elena’s job is not to say no; it is to write the conditions under which yes stays safe and to give every risk a named owner. Her instruments are a permitted-use rule that binds the model to the purpose actually reviewed, human review where the context lives, monitoring with a person’s name on it, least-privilege access, and an expiry date on predictions. If a rule cannot be used to refuse a powerful colleague, it is decoration.Spot the theatre
Read each control, decide whether it would change what actually happens, then tap the card.
Quick check
1. Why does Elena refuse "scores may be used for any commercial purpose approved by a manager"?
2. What breaks in module 12 if predictions are deleted nightly?
3. Elena's controls are handed to Kofi as conditions rather than advice. What does that change?