Risk: Why Could One Supplier Do This to Us?
Nobody at Calder was negligent, and that is precisely the problem. Ingrid Sørensen reconstructs how two sensible decisions built a six-year single point of failure, and brings the board four controls instead of the twenty it expects.
What you'll learn
- Reconstruct how a serious exposure gets built out of decisions that were each correct at the time
- Name a risk precisely enough that it buys the right control, including the tier-two dependency nobody mapped
- Design a control plan short enough to survive the return of normality
On Thursday of the fourth week, Ingrid Sørensen is doing something nobody has had time for since 3 March: reading the whole story in one sitting. Dev Anand’s note is on top, containing the sentence the company has been too busy to absorb — one control board, one supplier, one chip maker, in every unit of every line, for six years.
Below it is the running cost, still being tallied by Claire Beaumont: $85,000 of air freight, $44,000 of broker boards, $115,000 of overtime, $180,000 of Kestrel tooling, $250,000 for the redesign. And below that, the thing that makes this week unlike any other in Ingrid’s four years here — the board has asked her to explain how this was allowed to happen. They have never asked before.
An engineering finding and a spending tally go in; a short, costed, deliberately unambitious list of permanent changes comes out.
What lands on Ingrid’s desk
Dev’s finding is the substance: consolidating onto the CB-40 put one component, from one supplier, into all 80 ovens a week across three lines, and has done since 2020. The money is the argument. The attention is the opportunity — Eleanor Vance has put “structural review” on the board agenda, and for once nobody asks whether it can wait until after the quarter.
Everything else is fixed and not hers to reopen: Tomasz’s allocation, Grace’s build-ahead dates, the conversations Marcus and Siân have already had. Ingrid is not being asked to fix March. She is being asked the only question left — why was one supplier in Penang able to stop a factory in Rockford, and what has to be true for that never to be possible again.
What a risk manager actually does
The caricature is a spreadsheet coloured red, amber and green, presented quarterly to people who stopped reading it years ago. The real job is narrower: find the exposures nobody owns because they sit between two departments, describe them precisely enough that a control can be attached, and argue for money now against a loss that has not happened — the least persuasive argument in business, right up to the week it becomes the most persuasive.
Ingrid is one person with about a third of a quality engineer’s time. Calder buys roughly 1,400 parts from about 210 suppliers. Any plan assuming more capacity than that is a plan for a company she does not work at.
How risk people talk
- Inherent risk
- How bad something would be with no protection at all. Inherent, the CB-40 exposure is the whole factory.
- Residual risk
- What is left once your controls work as intended. If the control is a shelf of boards, the residual risk is how many weeks that shelf lasts.
- Control
- Something that actually happens — a stock level, a sign-off step, a report — rather than an intention. If nobody would notice it being skipped, it is not a control.
The software on Ingrid’s desk
Two queries would have found this exposure years ago. Nobody ran them.
The uncomfortable observation in this module is that the exposure was discoverable at any point in six years by joining two systems nobody thought to join. The parts database knows which components are shared across product lines. The ERP knows which of those come from a single supplier. Run those two questions together and the CB-40 appears at the top of a very short list.
The risk register is where the finding becomes a control with a named owner rather than a memo, and the discipline is that a register entry without an owner and evidence is decoration. The last screen is the one Ingrid cares most about: supplier concentration reported to the board quarterly, alongside revenue and safety. That single recurring number is what makes the next slow accumulation of sensible decisions visible while it is still cheap to undo.
The software on this desk
- PLM (Windchill)
- The parts database. Answers which components are shared across products — half of the query that finds a single point of failure.
- SAP spend analysis
- Which parts come from one supplier, and how much revenue depends on them. The other half.
- ServiceNow GRC
- The risk register: risk, control, owner, evidence. Without the last two it is a list of worries.
- Power BI
- The quarterly board number. Recurring visibility is what makes a control survive the return of normality.
The decisions
How the exposure was actually built
The temptation in week four of an expensive crisis is to find the decision that caused it. Ingrid looks, finds two, and reports that both were correct.
In 2020 Calder ran three control boards, one per line, each with its own firmware, spares holding and certification evidence. A project consolidated them into the CB-40: one part number at 80 a week instead of three at modest volumes, which cut the unit price materially, left stores holding one board rather than three, put one spare in every service van and replaced three sets of test evidence with one. It was written up as a good project because it was one.
Three years later a buyer found that splitting CB-40 volume across two board houses was costing money. Giving Vantor all of it moved the price from about $350 to $320 — on four thousand boards a year, roughly $125,000 annually, permanently, for a negotiation. It hit a savings target and went to the board as a win, which it was.
Nowhere does anyone write the sentence that mattered: that together the two decisions meant one factory in Penang could stop production in Rockford. Not because anyone hid it, but because they were three years apart and judged against different criteria — the consolidation as an engineering question, the sole award as a price one.
Risk accumulates in the gaps between good decisions, and it is the hardest kind to see because every step has a defender who is right: challenge the consolidation and an engineer shows you what three certified variants cost; challenge the sole award and a buyer shows you the saving in the accounts. Only the combination is dangerous, and combinations have no owner. Which is why Ingrid names nobody — a review that ends with a culprit tells everyone the mechanism has been dealt with, when nothing about it has.
Naming the risk precisely enough to act on it
The second decision is about language, which sounds like the soft part of this module and decides whether the money is spent well.
A single point of failure is a component, supplier, system or person whose loss stops everything, with no parallel path. Concentration risk is the wider condition of too much depending on one place — one supplier, one country, one customer, one engineer who understands the burner firmware. Tiering is how far up the chain you can see: tier one is who you buy from and hold contracts with, tier two is who they buy from. Most companies know the first and not the second.
Which is the refinement that makes this assessment worth reading. Calder’s risk was never really Vantor, whose factory is running and whose behaviour throughout has been honest. The binding constraint is the microcontroller inside the CB-40, allocated by a chip maker Calder has never contracted with, never spoken to, and which appears in no Calder system anywhere. The shortage came from two levels up, through a supplier as trapped by it as Rockford is.
That distinction has teeth. Name the risk “Vantor” and the obvious control is a second board house — but if Kestrel’s board used the same scarce microcontroller family, Calder would have paid $180,000 of tooling and 44% more per board for a second supplier that fails in the same week as the first. The control works only because Dev’s redesign specifies a microcontroller at least two manufacturers make in interchangeable form.
A badly named risk buys the wrong control
Mapping every sub-tier across 1,400 parts is a fantasy for a company this size. Asking one question about nine parts is not: what is the scarcest single thing inside this, and who makes it?What to change — and, more importantly, what not to
The third decision is the one Ingrid expects to defend hardest, because it involves proposing less than everyone wants.
After an incident the instinct is total: dual-source everything, hold six months of stock, audit every supplier quarterly. It is satisfying and unaffordable. Qualifying one second source for a safety-approved part cost ten weeks and $180,000; six months of cover ties up cash in a business about to be strained by this very crisis; quarterly audits of 210 suppliers is four audits a week, performed by one risk manager and a third of a quality engineer. And the board would probably approve all of it this month, which is the trap — within a year it is abandoned without anyone deciding to, the buffer consumed in the first tight quarter and never rebuilt, the audits slipping to half-yearly and then to “on exception”. A control skipped once with no consequence is a control that has finished.
So she proposes a filter of three conditions applied together. A part gets the expensive treatment only if it is single-sourced, and used across more than one line, and has a lead time longer than eight weeks. Any one alone is ordinary; all three together is what turns a supplier’s bad month into a stopped factory. Run against Calder’s parts list it returns nine components — the CB-40, a gas valve assembly from one Italian maker, the door-seal extrusion, a fan motor, the safety thermostat and four others. Nine, not nine hundred, out of 1,400.
For those nine, and only those nine: a qualified alternative or genuine second source, a defined buffer stated as weeks of cover, and a named owner. The rest of the base gets what it gets today, which now becomes a written acceptance of risk rather than an oversight.
Making it survive the return of normality
The fourth decision is the real lesson, and Ingrid puts it in one sentence: the risk register is not the control. A register records that somebody was worried — the CB-40 sat on Calder’s two years ago, rated amber, reviewed twice, noted both times, and changed nothing. The control is a small number of things that happen whether or not anyone remembers March.
The first is availability written into design sign-off, Dev’s point from module four made structural. No new design passes review if it depends on a single-manufacturer part without an explicit, named waiver. It costs nothing, it happens at the one moment when changing the answer is cheap, and it stops the next CB-40 being created rather than discovered.
The second is a supplier concentration figure in the quarterly board pack, beside the sales numbers and requiring no commentary: the share of weekly revenue passing through a component with no qualified alternative. Today that is 100%, because of the CB-40. A number that embarrassing, reported without anyone having to raise it, does more than any amount of advocacy.
The third is buffer levels agreed as policy rather than as a favour to operations. The 520 boards on the shelf in March existed because a planner once argued successfully for them, which means the next planner can argue them away; as policy, a buffer is a number in the system with an owner and a breach report. The fourth is one named owner per critical part — a person, not a function, because departments do not notice things and people do, but only when the thing is theirs.
That is four items, proposed on purpose to a board expecting fifteen. A long list is not a stronger response but a weaker one wearing more clothes: it becomes a programme, then a plan, then a status report where a third of the actions are closed, a third restated and a third quietly dropped.
The window is about six weeks
This conversation only ever happens after an incident. Ingrid’s analysis is essentially the one she could have written in January; what changed is not the evidence but that everyone has now felt it. Attention decays fast — at Calder about six weeks, ending roughly when deliveries resume. So she brings a good-enough paper in week four rather than a better one in week twelve, and asks the board to decide, not to note.Where this goes wrong
In most companies the post-incident review begins with “how did we let this happen” and ends three weeks later at the buyer who consolidated the volume — a satisfying finding, and conveniently that buyer has often moved on and cannot argue.
Three things follow. The organisation concludes the exposure has been dealt with because someone was held to account, while the gap between engineering and commercial decisions stays exactly as open as before. The successor learns that surfacing a trade-off is dangerous. And everyone else who knows of an awkward dependency in their own area now knows what raising it costs. The lasting damage is not the punished individual; it is the near miss nobody mentions for two years.
What Ingrid hands on
To Eleanor Vance and the board goes a paper in two halves. The assessment reconstructs how the exposure was built from two correct decisions, names it as a tier-two microcontroller dependency rather than a supplier problem, and lists the nine parts that meet all three conditions. The plan is four controls with prices attached: the CB-40 work is already committed through Kestrel and the redesign, so what is genuinely new is about $120,000 of qualification work on the other eight parts over eighteen months, roughly $200,000 of working capital held in defined buffers, and three governance changes costing almost nothing beyond the willingness to repeat them every quarter forever.
The constraint Eleanor inherits is deliberate. Once the concentration figure is in the board pack it cannot come out without a minuted decision to remove it, and once the plan is four items rather than twenty there is nowhere to hide a failure to deliver it.
The bottom line
Calder’s exposure was built by two individually correct decisions three years apart, in the gap where neither function’s review was looking — which is why blame teaches nothing and guarantees a repeat. Naming matters, because the real dependency was the microcontroller two tiers up, not the supplier. And the response is deliberately small: nine parts, four controls, priced — because a short list survives the return of normality and a long one does not.Spot the decision
Read each situation and decide how a risk manager should handle it, then tap a card to check.
Quick check
1. Why does Ingrid refuse to attribute the exposure to anyone's mistake?
2. What was the dependency Calder had never mapped?
3. Why does the control plan cover about nine parts rather than the whole supply base?