← AI Agents at Work: A Product Launch, End to End
Module 12 Free 11 min

Did It Work — And Were the Agents Worth It?

Lena Park and Arthur Reyes have $7.4m of pipeline against a $9m target and a ledger for the agents that nobody will enjoy. Reading an 82% result honestly, and measuring an agent by the whole system rather than the step that got faster.

What you'll learn

  • Read a missed target by finding which stage of the funnel underperformed, and why
  • Tell the difference between pipeline as a forecast and closed revenue as an achievement
  • Measure an agent across the whole system, including the verification effort its output created

Two quarters after launch, Lena Park and Arthur Reyes are in a room with the same spreadsheet open and slightly different jobs. Lena’s job is to say what happened. Arthur’s is to say what it was worth. The headline is that the campaign produced $7.4m of qualified pipeline against a target of $9m, which is 82%, and everyone who walks past the room already has an opinion about whether that is good.

The second question is newer and harder, and it is why Delia asked for this meeting rather than a slide. Four desks in this campaign ran AI agents. The tooling cost $38,000 — about three per cent of the budget — and the received wisdom in the building is that they were an enormous success, because everybody upstream experienced them as one. Arthur’s ledger tells a more complicated story, and it is the more useful one.

CURTIS & DANAoutcomes, sources, reasonsLENA & ARTHURdid it work, and was it worth it?DELIA & THE BOARDthe honest read, and what to fund

Twelve desks of decisions arrive here as four numbers — and the job is to say what they actually mean.

What lands on Lena’s and Arthur’s desks

A database that means one thing, because of two decisions taken three months earlier.

From Dana in module 10 comes a consistent tracking taxonomy, journey stages that exist as fields rather than as concepts, and one definition of MQL that held for the whole period. From Curtis in module 11 come the opportunities with their sources, and the far more analytically valuable half: every rejected MQL with a stated reason, from a closed list.

It is worth naming what those two hand-offs prevented, because it is invisible when it works. Without the taxonomy this analysis would have begun with a week of reconciling four spellings of LinkedIn. Without the signed definition, marketing and sales would have arrived with different MQL counts and the meeting would have been about whose number was real. Without the rejection reasons, Lena could tell you the campaign missed but not why — and a miss without a diagnosis is just a mood.

What analytics and finance actually do here

One establishes what happened. The other establishes what it was worth. They are not the same question and they are frequently confused.

Lena’s discipline is attribution and diagnosis: connecting spend to outcomes through a chain of records, then locating the stage where reality diverged from plan. Arthur’s is valuation: deciding what a result is worth against what it cost, including the costs that landed in somebody else’s budget. Marketing analytics without finance produces impressive percentages. Finance without analytics produces confident conclusions about the wrong thing.

The vocabulary of the wash-up

Conversion rate
The proportion moving from one funnel stage to the next. Only meaningful when both stages have a fixed definition for the whole period.
Pipeline
The value of open opportunities. A forecast of revenue that might arrive, weighted by nothing unless you weight it.
Closed revenue
Money actually contracted. The only figure in the report that has survived contact with a buyer’s signature.
Cost per opportunity
Total campaign spend divided by opportunities produced. The number a campaign is honestly judged on.
Customer acquisition cost
Spend divided by customers actually won. Falls over time as more of the pipeline closes, which is why quoting it early flatters or damns unfairly.
System boundary
What you decided to count. The single most manipulable choice in any efficiency claim, and the one nobody states.

The software behind the final number

Four screens, and the argument is always about which one is telling the truth.
Power BIChannel performanceLeads, MQLs, opportunitiesby source — only becausethe taxonomy held.SalesforcePipeline and closed$7.4m open, $1.9m closed.One is a forecast, one ismoney.ExcelThe agent ledgerTime saved against toolingand review. The number thatflatters, and the real one.PowerPointOne pageWhat worked, what did not,and what Cadence fundsnext.

Pipeline is a forecast. Closed revenue is a fact. Both belong on the page.

Reporting can attribute performance by channel only because Dana enforced one naming convention twelve weeks earlier; this screen is the payoff for the most boring decision in the course. The CRM supplies the two numbers that must appear together — open pipeline, which is a forecast, and closed revenue, which is not — because a campaign judged on pipeline alone can look excellent for two quarters and then quietly not convert.

The agent ledger is the module’s real contribution, and it is built in a spreadsheet because no tool computes it: time saved on one side, tooling and additional review on the other. Measuring only the step that got faster produces a spectacular and false number. And the board slide does what it has done in all three cases — reduces twelve people’s quarter to one page, which is honest only if somebody insisted on putting the uncomfortable column on it.

The software on this desk

Power BI
Channel performance, made possible entirely by a consistent tracking taxonomy set before launch.
Salesforce
Open pipeline and closed revenue. The first is a forecast; only the second is money.
Excel
The agent ledger — savings against tooling and added review. No tool produces this; someone has to build it.
PowerPoint
One page for the board, including the column nobody enjoys presenting.

Part one: did the campaign work?

Eighty-two per cent is not a verdict. It is the beginning of a question.

The headline, and why it is not yet an answer

The campaign spent $1.2m, exactly on budget, and produced 154 opportunities against a plan of 190, worth $7.4m of qualified pipeline against a target of $9m. The cost per opportunity came in at $7,800 against a plan of $6,300 — twenty-four per cent over.

That is the whole headline, and a great many organisations stop there and hold a meeting about whether the campaign was a success. Lena refuses the framing, and her reasoning is the most portable thing in this module: a result of 82% is neither success nor failure until you know which stage of the funnel underperformed and why. The same 82% can mean four completely different things. It can mean the audience was wrong, the message failed, the conversion machinery leaked, or the market was right and the timing was not. Those four diagnoses lead to four different decisions next quarter, and three of them would be actively harmful if applied to the wrong cause.

So she takes the funnel apart, stage by stage, against Delia’s original arithmetic from module 1.

Where the funnel actually broke

Leads: 2,690 against 2,530 planned. The campaign over-delivered by six per cent at the top, at a cost per lead of $446 against $474 planned. Whatever went wrong, it was not that nobody responded.

MQLs: 742 against 760 planned. Ninety-eight per cent — close enough to plan that it is noise. The lead-to-MQL conversion came in at 27.6% against 30% assumed, which is a small miss made invisible by the extra volume at the top.

Opportunities: 154 against 190 planned. Here is the break. MQL to opportunity ran at 20.8% against the 25% the whole plan was built on. The funnel held all the way down and then failed at the last join, which is a genuinely useful finding, because it rules out three of the four diagnoses immediately. The audience was reachable. The message worked. The plumbing worked. Something happened between an MQL and a salesperson accepting it.

This is where Curtis’s rejection reasons earn their existence. Of the roughly 588 MQLs that did not become opportunities, two categories dominate. Ninety-six were rejected as runs SAP and needs it now or timeline too short for a six-week implementation. And 121 were rejected as no budget this year — good companies at the wrong moment, sitting in a dated nurture track rather than in a bin.

Put a number on the first of those. Ninety-six MQLs, at the plan’s 25% conversion, is about 24 opportunities and roughly $1.15m of pipeline, and the connector that blocks them ships next quarter. That is not lost pipeline. It is deferred pipeline with an acquisition cost of zero, and it is the single most decision-relevant sentence in Lena’s analysis — which is only sayable because Wes published the limitation in module 2, Ryan put it in the evaluating stage in module 6, and Curtis made it a mandatory question in minute six of the first call in module 11.

The second finding is smaller and more embarrassing. Median first-touch time settled at 22 minutes in business hours from week four onward, but in the first three weeks of launch it was four and a half hours, because the duty rota was not staffed for a launch surge nobody had modelled. Leads contacted inside fifteen minutes became opportunities at nearly twice the rate of those contacted the next day. Lena estimates that cost around nine opportunities, and she is careful to call it an estimate.

Those two causes account for most of the 36-opportunity shortfall and not all of it, and she leaves the residual unexplained on purpose. An analysis that reconciles perfectly to the last unit has almost always had a story fitted to it.

The ranking rearranges again

Ryan’s lesson in module 6 was that cost per lead misleads and cost per MQL is truer. One stage later it happens again: LinkedIn looked like the best value in week six on cost per MQL, and its MQLs converted to opportunities at 14% against the campaign’s 21%, while the trade show and the webinars — which took Ryan’s reserve under his own rule — beat plan at the opportunity stage. Every metric flatters the stage at which it is measured. The only cure is to keep dividing one stage further down than feels necessary.

Pipeline is a forecast; closed revenue is the achievement

Arthur’s contribution begins with a distinction that gets blurred in almost every marketing review he has ever attended. Pipeline is not money. It is the value of deals that are open, which means the value of conversations that have not yet gone wrong. A $7.4m pipeline is a forecast made by the people most motivated to be optimistic about it, and treating it as an achievement is how organisations celebrate in March and discover in September.

The number he cares about is closed revenue after two quarters: $1.9m. About forty customers, at the $48,000 average contract. That is roughly a quarter of the opportunities created, and the rest are still open, still real, and still capable of disappearing.

Two things follow, and they point in opposite directions, which is why he insists on saying both.

The pessimistic read: $1.2m spent, $1.9m contracted. On a naive first-year view the campaign has barely paid for itself, and $1.2m divided by forty customers is $30,000 of acquisition cost each against a $48,000 first-year contract. Uncomfortable.

The optimistic read, and the more accurate one: that $30,000 is a ceiling that will fall, because the campaign’s cost is spread across every customer it eventually produces and $7.4m of pipeline is still open. And the revenue is recurring — the second year of each contract arrives with no acquisition cost attached. A $48,000 contract that renews twice is worth far more than $30,000 to win. Arthur’s judgement is that the campaign is sound and the target was missed, and that those two statements are compatible.

The most common misreport in marketing

Reporting pipeline as though it were revenue, then reporting the revenue separately six months later when nobody is looking. Pipeline answers “is there enough in front of us?”; closed revenue answers “did any of this work?” A report that gives only the first is not a result. It is an intention with a dollar sign on it.

Part two: were the agents worth it?

The flattering answer is available in about ninety seconds, and it is wrong by a factor of four.

The number that would have been reported

Here is the analysis Cadence would have produced if nobody had thought about it carefully. Content drafting used to take six weeks and took two. Market research took three weeks and took three days. Creative adaptation across forty placements took a fortnight of a junior designer and took a day. Total tooling: $38,000. Conclusion: agents cut production time by two thirds for three per cent of the budget, so expand everywhere.

Every sentence in that paragraph is true. The conclusion is still wrong, and the mechanism by which it goes wrong is worth understanding because it is nearly universal. Measuring only the step that got faster produces a wildly flattering number, for a structural reason: the saving and the cost land in different places. Production sits in the marketing budget, where the saving is visible and somebody is rewarded for it. Verification sits in legal, in product and in operations, where the cost is absorbed as workload rather than invoiced. No single manager sees both sides, so the organisation genuinely believes the flattering number. Nobody is lying.

The honest method is not complicated, and it is the whole methodological point of this case. Draw the system boundary around the entire process — from brief to approved, published asset — and measure it before and after. Not from brief to first draft. The draft was never the deliverable.

The ledger, told without hype

Arthur builds it that way, and it is not a triumphant document.

On the credit side, the agents saved roughly $180,000 of agency and analyst time. That is a real number with invoices behind it: research that would have been outsourced, copywriting that would have been commissioned, design adaptation that would have been billed by the hour. Anika’s research agent alone did 2,400 searches for $1,900 in three days, replacing about three weeks of analyst work.

On the debit side, tooling cost $38,000. And then the cost that nobody planned for: legal and factual review rose from about 40 hours a quarter to 130 hours. Fully loaded — Miriam’s time, outside counsel on the rights questions in module 9, building the claims register, Anika’s provenance discipline in module 3, Theo’s provenance logging in module 8, and Dana’s approved-asset library in module 10 — that is some $95,000 of extra verification effort.

So: $180,000 saved, $133,000 spent to save it. Net saving: about $47,000. Against a $1.2m campaign that is under four per cent, which is to say almost nothing. On money alone, the agents were roughly a wash.

Arthur says this out loud in the meeting, because the room expects a triumph and the room has been telling itself one for a quarter. Then he says the second thing, which is the one that matters.

What the agents were actually worth

The campaign reached market four weeks earlier than it could have without them. That is the return, and it dwarfs everything on the ledger.

Lena sizes it rather than gesturing at it. The campaign generated $7.4m of pipeline across roughly twenty-six weeks — about $285,000 of pipeline a week at the achieved run rate. Four weeks is therefore on the order of $1.1m of pipeline, which is more than twenty times the net cash saving. And she immediately qualifies it, because this is exactly the sort of number that gets quoted without its caveat: most of that is pipeline brought forward, not pipeline created. The honest claim is that Cadence spent a quarter of a year in market that it would otherwise have spent preparing, in a category where Arbor Risk is entrenched and Nodefield is moving quickly, and that being early is worth more in a launch than in almost any other kind of campaign.

Which produces the finding Delia takes to the board: the agents were not a cost-saving. They were a speed purchase, and it was worth making. The $38,000 bought time, and the $95,000 was the price of that time, paid in a different currency by different people.

Where the money actually went

Production got cheaper and checking got dearer, and the two nearly cancelled. If you take one number from this case, take that one — because the version circulating in most companies right now is the $180,000, on its own, with no mention of the 90 extra hours in legal.

Where this goes wrong

Two failure modes, and one of them is currently being celebrated in a great many quarterly reviews.

The first is the partial boundary. An efficiency claim is only as honest as the edge somebody drew around it, and the edge is almost never stated. “Content costs down 67%” is unanswerable unless you know whether review was inside the box. Ask of any productivity number: what was counted, over what period, and whose budget absorbed what fell outside. If nobody can answer, the number is a marketing claim about marketing.

The second is the opposite error and it is just as expensive. Having discovered that the net cash saving was $47,000, a finance team can conclude that the agents were not worth it and switch them off — and lose the four weeks, which was the entire return. Measuring the whole system is not the same as reducing everything to cash. The only reason Cadence can see the speed gain at all is that Arthur looked past the ledger he had just built.

What Cadence funds next

The recommendation is not “more agents” and not “fewer”. It is that the two budgets move together.

Delia, Lena and Arthur put five things to the board.

The research agent is funded again and expanded, because it is the clearest win in the case: $1,900, three days, provenance attached at collection, and 340 verified claims that made Miriam’s job faster rather than slower. It is the one agent in the campaign that reduced downstream verification work instead of creating it, and the reason is entirely Anika’s design decision in module 3.

The claims register becomes permanent infrastructure with a named owner, a budget and a calendar for its expiry dates. It is the highest-return item on the list, because it changes the shape of the cost curve: review effort stops scaling with volume of output and starts scaling with genuinely new claims. The register is also the reason the SAP limitation gets updated the week the connector ships rather than six months later.

Verification is budgeted as a line item scaled to planned output, not estimated from last year’s hours. This is the direct lesson of Ryan’s one-week estimate in module 6. Next quarter the agent tooling rises from $38,000 to about $60,000, and the review budget rises with it, in the same paper, approved at the same meeting. Cadence does not get to approve one without the other again.

The money moves to what converted: more into webinars and the trade-show programme, which won at the opportunity stage, and away from cheap-volume channels that won at the lead stage and lost everywhere after it.

And the 96 SAP-blocked MQLs get a dated campaign the week the connector ships — roughly $1.15m of pipeline at an acquisition cost that has already been paid.

The lesson of the collection

Three cases, three ways a decision travels — and one thing that cannot be handed to a machine.

Go back to module 1 and Delia’s arithmetic. Nine million dollars became 190 opportunities became 760 MQLs became 2,530 leads at $474 each, and every desk in the eleven modules after it either protected that ratio or quietly damaged it. Wes protected it by publishing the SAP gap, which cost leads at the top and saved opportunities at the bottom. Ryan damaged it slightly by planning a review week from a history that no longer applied. Miriam protected it by refusing a sentence that would have been indefensible in front of a customer. Curtis protected it by disqualifying people in minute six. The final cost per opportunity of $7,800 is the sum of those decisions, and none of them was made by an agent.

And go back to Delia’s other sentence, the one that looked like boilerplate in module 1: an agent may produce anything, and a named human owns every claim that reaches a customer. By module 9 it was the most consequential thing she wrote all quarter, and this analysis is what it cost — 90 extra hours in legal, $95,000 of verification, and a claims register that did not exist before. It was cheap at the price. The alternative was a launch campaign asserting something Cadence could not stand behind, to fourteen thousand companies, at the exact moment its reputation in a new category was being formed.

Which is the case’s lesson, and it generalises well beyond marketing. Agents move the work rather than removing it. Production becomes cheap; verification, approval and accountability become the bottleneck — and those are precisely the parts that cannot be delegated to an agent, because what they consist of is somebody being answerable. A model can draft the claim. It cannot be the person who is wrong if the claim is wrong.

Set the three cases side by side and the difference is instructive. In the first, a data project, decisions made early constrained the people downstream — a definition of churn written in week two determined what an operations team could act on in week twenty. In the second, a supply shock, decisions landed on customers and cash within weeks, and could not be taken back. In this one, the making got cheap and the checking became the job — which is the dynamic most organisations are living through right now, usually without having named it.

The uncomfortable implication is worth stating plainly. If your organisation adopts agents and its verification capacity does not grow, it has not become more productive. It has become faster at producing things nobody has checked, and it will find out which ones mattered in public.

The bottom line

An 82% result is neither success nor failure until you find the stage that broke: Cadence’s funnel held to MQL and failed at 20.8% MQL-to-opportunity, with 96 SAP-blocked leads worth about $1.15m of deferred pipeline. Pipeline is a forecast; $1.9m of closed revenue is the achievement. On the agents, the honest ledger is $38,000 of tooling plus $95,000 of extra verification against $180,000 saved — a net $47,000, almost nothing on a $1.2m campaign; the real return was four weeks earlier to market. Measure the whole system, not the step that got faster — because agents move the work rather than removing it, and the work they move it to is the part nobody can delegate.

Designing this desk’s agent: the reporting agent

An agent that assembles the numbers, and must never choose which story they tell.

Lena and Arthur’s agent does the assembly — pulling from four systems into one comparable picture. What it must not do is the part that makes reporting honest: deciding what to include when the answer is unflattering.

What this agent actually is

State it needs
The attribution model in force, the reporting period, and which caveats are outstanding.
Inputs
The data warehouse, the CRM, ad platform APIs, and the agent cost ledger.
Core behaviours
Aggregate by channel and stage, apply the chosen attribution model, and draft a narrative summary.
Constraints — what it may not do alone
It may not choose the attribution model, may not report pipeline without closed revenue beside it, may not exclude a channel or period, and may not publish a summary with an empty caveats list.

One concrete design choice. Make caveats a required, non-empty field. A report that cannot be saved without naming what would change the conclusion is a report that keeps its authors honest — including about the agent ledger itself, where the saving and the cost land in different budgets.

{
  "period": "2026-Q3",
  "channel": "all",
  "leads": 2690,
  "mqls": 742,
  "opportunities": 154,
  "pipeline_usd": 7392000,
  "closed_usd": 1900000,
  "spend_usd": 1200000,
  "attribution_model": "first_touch",
  "caveats": ["96 rejections were timing-related, not fit", "closed figure covers two quarters only"]
}

The metric to track. Forecast calibration — of the pipeline reported as likely to close, what share actually closed? Run that comparison for two quarters and the word pipeline stops being a trophy and becomes a forecast with a known error rate, which is the only form in which it is useful.

Failure modes and moral hazards

Measuring only what got faster: count the drafting time saved, ignore the review time created, and produce a spectacular return that is false. Attribution chosen to flatter: first-touch and last-touch tell different stories, and picking after seeing both is how a campaign proves whatever it wanted to. Confident narrative, thin sample: a fluent explanation of why one channel outperformed, generated from a difference that is within noise.

Human responsibility statement

Lena and Arthur own the number the board believes. Between them they decide what goes on one page, and the temptation they are managing is not fraud but selection — the honest report is the one that includes the column nobody enjoys presenting.

Read the number honestly

Read each finding and decide what it actually tells you, then tap a card to check.

Quick check

1. Where did Cadence's funnel actually underperform?

2. Why does measuring only the step that got faster flatter the agents?

3. What was the real return on the agents at Cadence?