← AI Agents at Work: A Product Launch, End to End
Module 4 Free 9 min

Customer Insights: What Do They Actually Feel?

Joel Brennan takes 340 sourced claims about what the market says and sets them against twelve interviews, a year of CRM notes and 2,100 support tickets — and finds that the market talks about dashboards and buys because someone was blamed in front of their boss.

What you'll learn

  • Explain why published evidence systematically under-represents the reasons people actually buy
  • Triangulate four sources of customer evidence and judge what each one is good and bad for
  • Grade a finding by confidence, and say plainly what your evidence cannot tell you

On Wednesday morning Joel Brennan has Anika’s file open — 340 claims, each with a URL, a date and a source tier — and beside it a list of twelve people who have agreed to talk to him for forty-five minutes each. He has three and a half days before Camille Duarte needs something she can build positioning on, and the entire value of those three and a half days lies in the difference between the two things on his desk.

The file tells him what the supply chain risk market says about itself. It says a great deal, fluently, with citations. What it does not contain is a single sentence written by somebody who was not trying to be read.

ANIKA RAGHUNATHANwhat the market saysJOEL BRENNANwhat do they actually feel?CAMILLE DUARTE, PMMneeds, graded by confidence

Breadth arrives from the left; depth is added here — and the gap between them is where the campaign's message comes from.

What lands on Joel’s desk

A dataset with its evidence attached, and a warning written across the top of it.

Anika’s hand-off is unusually honest for a research deliverable. Every claim carries its source, its date and its tier — filing, trade publication, vendor material, review site, forum — so Joel can see not just what was said but what sort of thing said it. He also gets a written list of the questions the agent searched for and could not answer.

And he gets her standing warning, which decides how he spends his week: this dataset describes what the market says, not what customers feel. Joel’s job is to find out what is true of the people who published nothing.

The constraints above him are fixed. Delia drew the audience boundary in module 1 — mid-market discrete manufacturers in North America, $50m to $500m of revenue. Wes settled in module 2 what Supply Signal actually does: it flags nine risk signals earlier than a person would notice them, and it does not predict. Joel is not allowed to discover a customer need the product cannot meet and hand it on as an opportunity.

What a customer insights analyst actually does

Everyone in the building has an opinion about customers. This desk exists to make some of those opinions checkable.

Joel does not run focus groups and he does not own the survey tool. His job is to combine evidence of different kinds and different quality into a picture of what customers need, and — the part that makes him useful rather than decorative — to say how confident anyone should be in each part of it.

The distinction he lives by is between what people say and why they buy. These are rarely the same sentence, and the second is almost never spoken first. A buyer will tell you they need better visibility of supplier risk. Six questions later it emerges that what they need is to never again be the person in the Monday production meeting who did not know.

The vocabulary of customer insight

Triangulation
Checking a finding against several independent kinds of evidence. Not “more data” — differently-biased data, so the biases do not line up.
Selection bias
When the people you hear from differ systematically from the people you care about. Reviewers, forum posters and interviewees all volunteered; most of the market did not.
Confidence level
A grade written beside a finding saying how much weight it will bear. High, moderate or low, with the reason attached.
Verbatim
A customer’s exact words, kept intact. The raw material of a message, and the thing summaries destroy first.
Non-customer blind spot
Everything you cannot learn from your own customers, because they are the ones who said yes.

The software on Joel’s desk

Three systems hold what customers said. None of them holds how they felt.
SalesforceWhat sales heardCall notes, filtered throughwhat the rep hoped thecustomer meant.ZendeskWhat customers said2,100 tickets. Unflattering,unfiltered, and the bestsource in the building.DovetailThe interviewsTwelve transcripts, tagged.Depth, small numbers, andpolite answers.Clustering agentFind what to readGroups 2,100 tickets intothemes — including one ofninety he'd have missed.

The agent finds the ninety tickets worth reading. A person still has to read them.

CRM notes record what sales heard, which is not the same as what the customer said; every note has passed through a person who wanted the deal. Support tickets have no such filter, which is exactly why they are the most valuable text in the company — nobody writes a support ticket to be polite. Interview transcripts in a research tool supply depth and cannot supply numbers.

The fourth screen is the agent, and its role here is the most instructive in the course. It clusters two thousand tickets into themes in minutes, which is genuinely impossible by hand, and surfaces a group of ninety Joel would never have sampled his way into. Then its summary describes them as dissatisfaction with notification timeliness — accurate, useless, and flattened. He reads forty of the raw ones and finds humiliation. Use the agent to find the ninety; read the ninety yourself.

The software on this desk

Salesforce notes
What sales heard, recorded by someone who wanted the deal. Useful, and never neutral.
Zendesk
Support tickets: unfiltered, unflattering, and the closest thing to a transcript of the relationship.
Dovetail
Interview transcripts, tagged and searchable. Depth without numbers.
A clustering agent
Groups thousands of tickets into themes so a human knows which forty to read. It finds; it does not feel.

What Joel works out this week

One of these findings changes the campaign. The rest are what make it safe to believe.

Why the agent’s dataset is not enough

The research agent read the public internet thoroughly — 2,400 searches, more sources than any analyst would reach in three weeks. The limitation is not effort. It is that the public internet is not a sample of buyers. It is a sample of people who wrote something down.

Consider who writes. Vendors write constantly, and every word is positioning. Analysts write, and their frame is the category, because the category is the product they sell. Journalists write about what changed. And a small, articulate minority of practitioners write: the ones with a conference talk to give, a consultancy to build, or a grievance loud enough to justify a review. All four groups are real; none is representative.

Now consider what nobody writes. Nobody blogs about the boring reason they bought — the incumbent contract was up and switching seemed manageable. Nobody publishes the embarrassing reason — they had already told the board it was handled. And nobody writes the unspoken one, which is usually about status: who will look competent, who will look careless, and in front of whom.

That is the systematic bias, and it is not random noise you can average away. Published material over-represents what people are willing to say in public and under-represents what actually drives a purchase — and the gap is widest where the emotion is strongest. Anika’s dataset is not wrong. It is a faithful record of a conversation held in public, by people conscious of an audience.

The four sources, and why he needs all of them

Joel works with four kinds of evidence, and he chooses them for how differently they fail.

Interviews give him depth and nothing else. In forty-five minutes he can follow a vague answer down four levels until it reaches something specific and human. The costs are that twelve conversations are twelve conversations — no proportion can be derived from them — and that an interviewee, being a polite professional, tells you what they think you want to hear. Half of Joel’s craft is asking about the last time something happened rather than what usually happens, because people invent their policies and remember their Tuesdays.

CRM notes are what sales heard: a genuine record of hundreds of conversations he was not in. They are also filtered twice — by what the buyer chose to say to someone trying to sell them something, and by what the salesperson wrote down, which is disproportionately whatever supported the deal progressing. A lost-deal note saying “budget” is true about as often as it is convenient.

Support tickets are the best source in the building and almost nobody in marketing reads them. They are unfiltered, unflattering and unperformed — nobody writing a ticket at ten past five is managing their personal brand. Their weakness is structural: everyone who files a ticket already bought. They tell you nothing about the people who evaluated Cadence and went elsewhere, or the far larger group who looked at the problem and did nothing.

Anika’s web evidence supplies what the other three cannot: breadth and currency. It covers the whole category, including competitors nobody at Cadence has met, and it is dated. What it has no access to is feeling.

Triangulation is the discipline of setting these against each other. Because each source is biased in a different direction, agreement between them means something. Joel’s working rule is blunt: a finding that appears in three of the four sources is worth acting on; a finding that appears in one is worth investigating. He writes both down. He treats them differently.

How he used an agent himself, and where he stopped trusting it

There are 2,100 support tickets across Cadence’s existing manufacturing products from the past eighteen months, and Joel cannot read them in three and a half days. So he used the approved internal tooling to cluster them by theme and summarise each cluster — with identifying details stripped first, because customer material going into a general-purpose tool is a disclosure of someone else’s confidential information. That is Miriam’s territory in module 9, and easier to get right on Wednesday than to explain in week eleven.

The clustering was excellent. It surfaced a group of roughly ninety tickets Joel would never have found by sampling, all circling the same shape of problem: a supplier issue discovered by the customer’s customer before it was discovered internally. The summary of that cluster read: users report dissatisfaction with the timeliness of supplier issue notification.

That sentence is a fair précis and it is worth nothing. So Joel read forty of the ninety raw tickets, and what was in them was not dissatisfaction. It was humiliation. “I found out from my customer.” “I had to stand in front of the plant manager and say I didn’t know.” “Second time I’ve looked like an idiot for something that was in the system.” The agent had done what summarisation does — averaged the language toward the neutral — and in averaging it had removed the only property of those tickets that mattered.

What summaries flatten

An agent is a superb instrument for finding where to look in 2,100 documents and a poor one for telling you what you will feel when you get there. Use it to locate the ninety. Then read the ninety.

The disagreement he finds

By Thursday Joel has a contradiction, and it is the most valuable thing produced by any desk this month.

The public conversation in Anika’s dataset — vendor pages, analyst notes, trade coverage — is about risk visibility. Dashboards, supplier scorecards, coverage of tiers two and three, a single pane of glass. Arbor Risk sells it that way, Nodefield sells a cheaper version of the same sentence, and every article about the category is organised around the same nouns.

The interviews and the tickets are about something much narrower and much less comfortable. Eight of the twelve interviewees, unprompted, described a specific incident in which a supplier problem became visible to them too late — and in every one of those accounts the sharpest detail was not the operational damage. It was who else found out first, and how it felt. A plant operations director described being asked in a Monday meeting why a line was down and having no answer, then said the sentence Joel writes at the top of his report: I never want to be surprised in front of my boss again. The ninety tickets say the same thing in the same register.

That contradiction generalises. A market talks in categories and buys for reasons. Categories are respectable, comparable and safe to be quoted on. Reasons are personal, occasionally undignified, and enormously predictive of whether someone signs. The published conversation will always be the first kind, because publishing is a public act.

The gap between the two is where positioning comes from. Anyone can write the category sentence; Arbor Risk already has, with more budget behind it. Only someone who has read the ninety tickets can write to the fear. Joel does not write the campaign message — that is Camille’s decision in module 5 and Naomi’s execution in module 7 — but he hands over the raw material and marks it, because a message built on the category competes on a field the incumbent owns, and a message built on the reason competes on ground nobody has claimed.

What he refuses to conclude

Joel’s last decision is about restraint, and it is what makes the rest of the report usable.

Twelve interviews cannot size a market. Eight of twelve is not 67% of anything; it is eight people, recruited through channels that favoured people willing to talk to a vendor. Joel will not let that number leave his desk as a percentage, for the same reason Anika refused to turn eleven of forty reviews into 27%: invented precision is harder to dislodge than an honest theme, because it looks as though it has been measured.

Support tickets cannot tell you about people who never became customers — the non-customer blind spot — and here that is most of the addressable market. The real competitor in this category is not Arbor Risk or Nodefield; it is the spreadsheet, and everyone using one is invisible to every internal source Joel has.

So he grades. Beside every finding sits a confidence level and the reason for it: high where three or four sources agree and at least one is external; moderate where two agree, or one strong source carries it alone; low where it rests on a single thread of evidence and is offered as a hypothesis to test rather than a fact to build on. The blame finding is high — interviews, tickets and forum language carry it from three differently-biased directions. That mid-market buyers will pay a premium for faster deployment is moderate. Anything about distributors and wholesalers is low, and labelled as such, because Joel interviewed nobody in that segment.

The grading is not modesty. A report that asserts everything with equal force is destroyed by its first challenge: when one finding turns out to be thin, the challenger has been handed permission to doubt all of them. A report that has graded itself survives, because the weak findings were already labelled weak and the strong ones are still standing when the argument ends.

Where this goes wrong

The common failure is not missing the insight. It is finding it and then losing it on the way to the slide.

In real companies the interview happens, the sentence gets said, somebody writes it down — and then the finding is rounded back up into the category before it reaches anyone with a budget. “Supply chain risk visibility” sounds like a market. “Never be surprised in front of your boss again” sounds unprofessional, and the person presenting it can feel the room deciding whether they are serious. So the specific becomes the general, the report reads like every other report, and the campaign competes on the incumbent’s ground with a fraction of the incumbent’s money.

The mirror-image failure is a report that asserts everything at the same volume. Somebody senior challenges the thinnest claim, wins, and the good finding goes down with it.

What Joel hands on

A customer needs report, every finding carrying its confidence, and one insight flagged as the one to build on.

Camille Duarte receives a document organised by need rather than by source, with each finding stating its confidence level, the sources behind it and, where it matters, the customer’s exact words. The verbatims travel intact — they are what Naomi will need in module 7, and no paraphrase of “I looked like an idiot” is worth having.

One finding is flagged above the others: the buyer’s real fear is being blamed for a shortage nobody saw coming, and their real want is never to be surprised in front of their boss again. High confidence, three independent sources, and directly served by what Wes says the product actually does — flagging signals earlier, which is a promise about warning rather than prediction.

The constraint Camille inherits is the grading itself. She may build the positioning on the high-confidence findings without further work. Anything marked moderate she may use but must not lead with. Anything marked low she must test before it becomes a promise — and since Joel’s evidence on distributors and wholesalers is entirely low-confidence, choosing that segment in module 5 would mean choosing it on public data alone, with no idea what those buyers feel.

The bottom line

An agent can read everything published and still miss the reason people buy, because published material records what people will say in public. Joel triangulates four differently-biased sources — interviews, CRM notes, support tickets and web evidence — and acts on what appears in three of four. The market talks about risk visibility; the buyer wants never to be surprised in front of their boss again, and the gap between those two sentences is the campaign’s positioning. Every finding leaves his desk with a confidence level attached, because a report that grades itself survives challenge.

Designing this desk’s agent: the clustering agent

An agent that decides what a human should read, and must never be the one who reads it.

Two thousand tickets is beyond a person and trivial for a clustering agent. The design question is not whether to use one — it is how to stop its summary becoming the finding.

What this agent actually is

State it needs
The clusters formed so far, their sizes, and which have been reviewed by a named person.
Inputs
Support tickets, CRM call notes and interview transcripts, all inside approved tooling with identifiers stripped before processing.
Core behaviours
Cluster by theme, size each cluster, summarise it, and surface exemplar items.
Constraints — what it may not do alone
It may not quote a named customer externally, may not infer anything about a person, may not size a market, and may not be the sole reader of any cluster it reports.

One concrete design choice. Make exemplar identifiers a required output and human review a required field. A cluster is not reportable until someone has opened at least five raw items and signed the row — which is what forces Joel to read the forty tickets where the agent said dissatisfaction and the customers said something much sharper.

{
  "cluster_id": "cl-14",
  "size": 90,
  "theme": "notified after the customer already knew",
  "exemplar_ids": ["ZD-40188", "ZD-40312", "ZD-40447"],
  "reviewed_by": "joel.brennan",
  "raw_items_read": 40,
  "confidence": "high"
}

The metric to track. Human-review coverage — the share of reported clusters where somebody actually read the raw material. Below 100% it is not an insight process, it is a summarisation process wearing one. Pair it with agreement between two readers on what a cluster means, which is how you find out whether the theme is in the data or in the reader.

Failure modes and moral hazards

Emotional flattening: humiliation is summarised as dissatisfaction, and the strongest finding in the dataset is neutralised into a phrase nobody would act on. Articulacy bias: clusters form around customers who write clearly and at length, which is not the same population as customers who are angry. Re-identification: the name is stripped and the detail that makes the company obvious is not — a privacy failure that passes every automated check.

Human responsibility statement

Joel owns the sentence “this is what customers feel.” It is one of the few claims in a company that cannot be traced to a system, which is precisely why a named person has to stand behind it.

How much weight will it bear?

Read each finding and decide what confidence you would put beside it, then tap a card.

Quick check

1. Why is Anika's 340-claim dataset insufficient on its own?

2. What makes triangulation across four sources worth the effort?

3. Why does Joel write a confidence level beside every finding?