Marketing — the Atlas · ch.10 · research & insight
📣 Chapter 10 · Part II · Strategy

People can't tell you why they buy

Every strategy in this book runs on knowing what buyers actually do. Ask them, and they'll tell you — honestly, helpfully, and wrong. This chapter is how to find out anyway.

Here's the whole chapter in one line: buyers are an unreliable narrator of their own buying — so watch what they do, mine what they say for language, and never bet a budget on a single stream of evidence. Everything below is that idea: the gap, the tools, the most expensive wrong question in history, and a pipeline that survives all of it.

1The reporting gap

Start with the uncomfortable data point every researcher learns first. Ask people whether they'll exercise three times a week, and most say yes; count who shows up, and most don't. Ask whether they'd pay more for the sustainable option, and around two-thirds nod; watch the checkout, and the premium-payers thin to a fraction. Ask "would you buy this at $49?" and a room full of sincere yeses converts into a launch that lands at a fifth of the forecast. Nobody was lying. Everybody was wrong.

Chapter 2 explained why. System 1 buys; System 2 answers the survey. The fast, associative machinery that actually chooses in the aisle is not the machinery that fills in questionnaires — that's the press secretary, and the press secretary reports the person you believe yourself to be: the one who exercises, reads contracts, and rewards virtuous brands. The buyer and the respondent share a body and not much else.

Programmer's version: asking users what they want is polling the cache, not the source. The answer comes back instantly and with total confidence, because it was precomputed from self-image and social expectation — and it was never once invalidated against actual behavior. The source of truth is the purchase log, the visit log, the churn table. The cache is the interview.

So research's first law, the one every method below gets graded against: watch beats ask. Not "never ask" — asking is how you learn the customer's words, and words are ammunition for everything in Chapter 1. But when watching and asking disagree, watching wins, every time, and they disagree constantly. Here's how big the gap runs.

Interactive · the say–do gap Click a row to see the gap and the mechanism behind it
claimed in the survey observed behavior
Click a row.
📈
Evidence check. The intention–behavior gap isn't folklore — it's one of the most replicated results in behavioral science. Meta-analyses of intention studies find that even a large change in stated intention produces only a small-to-medium change in actual behavior, and stated intentions explain well under a third of the variance in what people go on to do. In pricing research the effect has its own name — hypothetical bias — and the standard finding is that stated willingness-to-pay runs two to three times higher than what the same people pay when the money is real.

2The classic toolkit, honestly rated

Before David Ogilvy was advertising's most quotable man, he was a door-to-door researcher for George Gallup, and it shows in his most famous line: "The consumer is not a moron — she's your wife." The point wasn't chivalry. It was that guessing about customers from a Manhattan office is contempt, and research is respect. The classic toolkit he inherited and extended is still the toolkit — so let's rate it honestly, tool by tool.

Focus groups: great for language, terrible for prediction. Put eight strangers in a room with a moderator and you'll hear how real people phrase the problem, watch faces react in real time, and collect metaphors no one on your team would have invented. What you will not get is a forecast. Groups perform consensus; the loudest voice captures the room within minutes; and everyone answers as their imagined best self, out loud, in front of witnesses — the say–do gap with an audience. Mine focus groups for vocabulary and reactions. Never count them.

Surveys: good for sizing what you already understand. Ask about things people can accurately report — what they bought last week, which brands they've heard of, what they currently pay — and a well-sampled survey is a fine measuring stick. Ask hypotheticals — would you switch, would you pay, would you recommend — and you're back in the cache. A survey is an API whose responses are made up on the spot: it always returns 200 and a well-formed answer, and nothing in the payload tells you it was hallucinated at request time.

Panels: the gold standard behavioral record. A consumer panel — the same households' actual purchases, logged for years — is the closest thing marketing has to production telemetry. It's slow, unglamorous, and it's the dataset the Ehrenberg-Bass laws in Chapter 5 were mined from: double jeopardy, light-buyer growth, the lot. When this book says "the buying data insists," panels are the data doing the insisting.

And then there's the strangest branch of the classic era: Ernest Dichter's motivational research, which put products on the Freudian couch in the 1950s and came back with claims like "a convertible is a mistress." Brilliant, generative, occasionally right — and built so that no experiment could ever prove it wrong. Hold that thought.

"A convertible is a mistress. The sedan is a wife."
Ernest Dichter · Institute for Motivational Research
the methodDepth interviews on the analyst's couch: what does the product mean, underneath?
the insightConvertibles pulled men into showrooms; sensible sedans drove home. Dichter read it as desire vs. duty.
the catchBrilliant and unfalsifiable — no possible study could prove the mistress wrong.
the legacyMarketing learned buyers have motives they can't report. It just couldn't yet check which ones.
"The market is talking whether you ask or not."
search · reviews · communities · support logs · 2026
the methodContinuously mine what people type when nobody's grading them: queries, reviews, threads, tickets.
the upgradeFalsifiable at last — a claimed motive either shows up in the logs or it doesn't.
the catchEvery stream lies in a named direction: search sees problems, reviews sample extremes, communities skew devoted.
the ruleUse the streams — and say out loud which way each one is biased.

3New Coke, the canonical autopsy

April 23, 1985. Coca-Cola, rattled by years of losing blind sip tests to sweeter Pepsi, replaces its 99-year-old formula. This was not a hunch — it was the largest research program in the company's history: roughly 200,000 taste tests, run competently, at enormous cost. The sweeter new formula won them. Then the public got the news, and the company spent seventy-nine days under siege — thousands of furious calls a day, protest groups, people hoarding cases — until "Coca-Cola Classic" came back to a hero's welcome that July.

The autopsy matters because of what didn't go wrong. The sample wasn't too small; two hundred thousand is a census by research standards. Respondents didn't lie; they genuinely preferred the sweeter sip. The statistics held. The research answered its question correctly — and it was the wrong question. A three-ounce blind sip measures the liquid. The purchase was never about the liquid. It was about the brand from Chapter 6: identity, ritual, the red can at every childhood picnic, the sense that Coke belonged to its drinkers and not to Atlanta. Sip preference and buying loyalty are different variables, and only one of them was measured 200,000 times.

The bitterest detail: the warning was in the data. In some sessions where people were told the new formula would replace the old one — not join it — a vocal minority reacted with something close to anger. That signal was logged and set aside as noise, because it didn't fit the question being asked.

What would a right-question version have looked like? It would have tested the replacement, not the recipe: "the Coke you grew up with will no longer exist — how do you feel?" Full cans at home, over weeks, inside the ritual, with the loss made explicit. Same budget, same rigor, opposite finding. The method was never the problem. The aim was.

⚠️
The most dangerous research artifact is a precise answer to the wrong question. Precision reads as truth — 200,000 data points feel like armor — and so the wrongness travels straight to the boardroom wearing a lab coat. Before you fund a study, spend an hour on the question. Rigor is cheap to add later; aim isn't.

4Qual finds the question, quant sizes it

So the toolkit has two halves, and the whole craft is keeping them in the right order. Qualitative work — interviews, ethnography, standing in a parking lot at dawn — generates hypotheses and vocabulary. It's how McDonald's learned in Chapter 4 that a milkshake competes with bananas and boredom: not from a questionnaire, but from watching who bought and asking what job they were hiring for. Qual is an existence proof machine. It tells you a thing is real and hands you the customer's own words for it — which is exactly the fuel Chapter 1's positioning checklist demands ("words your customer wouldn't say out loud" is a research failure, not a writing failure).

Quantitative work — sizing surveys and controlled experiments — validates and measures. This creed is a century old: Claude Hopkins was running keyed coupons in 1923 and wrote in Scientific Advertising that almost any question can be answered "cheaply, quickly and finally, by a test campaign." Chapter 28 turns that into a full experimentation practice; for now the point is the division of labor: qual tells you a thing exists, quant tells you how much of it there is and whether it's really there.

Reverse the order and you manufacture confident nonsense. Run the big survey first and you get precise measurements of categories you invented at your desk — the respondents dutifully sort themselves into boxes nobody lives in. Run the focus group after the numbers and it becomes a cherry-picking machine: eight strangers will happily "explain" any statistic you show them. The pipeline runs one way: watch and listen until a sharp question exists, then size it.

And when you do size it, respect the noise. Small samples don't politely say "not enough data" — they hand you crisp, impressive-looking numbers that happen to be static. Run the machine below a few times with the true effect pinned to zero and watch how often n=15 delivers a "winner."

Interactive · the noise machine run 00
Each click is a fresh simulated test
Pick a sample size to run a simulated concept test. Then set the true effect to zero and keep running — count how many "results" the void produces.
💡
Why qual gets a pass on sample size. "You only talked to nine people" is a category error. Qual's job is existence proofs and vocabulary, and an existence proof needs one clean specimen, not a confidence interval. The sin isn't interviewing nine people — it's quoting them as percentages.

5Watching at scale

"Watch beats ask" used to mean standing in stores with a clipboard. Now the watching comes to you, in five streams — each honest in a way surveys can't be, and each biased in a way you're required to name before using it.

Behavioral analytics. What users actually click, abandon, repeat, and churn from — production telemetry for demand itself. Bias: it only sees people who already found you. The whole non-customer universe, where Chapter 5 says your growth lives, is invisible to your own dashboard.

Search data — the world's most honest survey. People confess to search boxes things they'd never tell an interviewer: the real symptom, the embarrassing comparison, the "is it normal that…". Nobody performs virtue for an empty text field. Bias: search sees problems, not satisfactions. Happy customers don't google; the stream is a river of friction with the delight filtered out.

Review mining. Thousands of customers explaining, unprompted and in their own vocabulary, what the product was hired for and where it failed — the three-star reviews are the seam, written by people with no axe to grind. Bias: reviews oversample the extremes. The distribution is J-shaped — the furious and the delighted write; the satisfied middle stays silent.

Support tickets and social listening. Nobody files a ticket to be polite, which makes ticket text some of the most honest language you own — and a spike is an early-warning siren. Social listening adds the unprompted conversation around the category. Bias: both sample the vocal. The loudest one percent generates half the text, and the loudest one percent is not your market.

Notice the pattern: none of these streams is a verdict. Each is a biased sensor, valuable precisely because its bias differs from the others'. That's the setup for Section 8 — but first, who's doing the watching, and how often?

📈
Evidence check. Seth Stephens-Davidowitz's Everybody Lies made the case with receipts: on topic after topic — prejudice, health, anxiety, actual intentions — aggregate search behavior diverges sharply from what the same populations tell surveys, and the search data tracks real-world outcomes better. The instrument works because no one is watching — which is also its limit: the moment you ask, the performance resumes.

6Continuous discovery

The classic failure shape of research isn't a bad study — it's a good study, once, three years ago. The big annual project ships a beautiful deck, the deck gets presented, and the knowledge starts decaying the day the last slide lands. Markets move; the deck doesn't. Batch-mode research in a streaming world.

The modern correction is continuous discovery, and its clearest articulation is Teresa Torres's: the people actually building the product hold weekly small customer touchpoints, as a habit rather than a project. Not a research department's quarterly report — the product trio itself, talking to a couple of customers every week, forever. The unit of research stops being "the study" and becomes "the cadence." Interviewing becomes infrastructure, like CI: nobody schedules a special initiative to run the tests; they just run.

The workhorse interview on that cadence is one this book has already met: the switch interview from Chapter 4. "Walk me back to the day you decided to sign up — what did you fire to hire us?" Run it with recent sign-ups and recent churns on a permanent loop and you get a live feed of the jobs you're being hired for, the real alternatives you're beating, and the moment demand actually forms — the milkshake study, running as a service instead of a story.

The other half of the correction is memory. Findings that live in slide decks evaporate — a deck is write-only memory, and six months later someone re-funds the same study because nobody can find the last one. An insight repository — every finding written down, tagged by source and evidence strength, searchable by anyone — is what lets discovery compound. Interviews append to it, review mining appends to it, experiments append to it; strategy reads from it. It is the difference between an organization that learns and one that repeatedly pays to be told.

💡
The cadence test. One question diagnoses a company's research health: "When did someone who builds the product last talk to a customer?" If the answer is a date, you have discovery. If the answer is the name of a project, you have archaeology.

7Evidence check — the synthetic respondent

The loudest research debate of 2025–26: if large language models can imitate people, why pay people? Prompt a model into a thousand personas — "you are a 34-year-old nurse in Leeds, price-sensitive, two kids" — and survey the lot in an afternoon for pennies. The pitch writes itself: infinite panel, zero recruiting, no incentives budget. And the results come back fast, fluent, and formatted. The question is whether they come back true.

The core failure mode has a name: distribution collapse. Human answers are messy — long tails, contradictions, the weird 4% who use the product in a way nobody designed for. Synthetic answers regress to the plausible middle: each persona drifts toward the most statistically likely response for its demographic sketch, so a thousand personas give you the variance of about a dozen. The mean can land close to a human panel's mean while the shape is wrong everywhere — and insight lives in the shape. The morning milkshake commuter was a tail. A synthetic panel would have smoothed him away.

Two more failure modes stack on top. Sycophancy: models lean toward agreeing with the premise of the question, which turns every concept test into applause. And training-data echo: a synthetic respondent is a compression of what the internet already said — polling it is polling the cache again, now with the cache dressed up as a thousand fresh users. It can tell you what people like your customers used to say. It cannot tell you what your market will do next, because nothing about your market is in the weights.

None of which makes the tool useless — it makes it a drafting tool. Use synthetic respondents to pressure-test questionnaire wording before spending real sample, to generate candidate segmentation hypotheses worth checking, to rehearse an interview guide against a plausible skeptic. Cheap, fast, genuinely helpful. Then stop. The line is bright: the moment a number would move a budget — demand estimates, willingness-to-pay, purchase intent, forecasted share — synthetic answers are inadmissible.

📈
Evidence check. The "silicon sampling" literature that piled up through 2024–26 converges on a consistent pattern: on coarse aggregates, LLM panels often correlate respectably with human polls — which is exactly what makes them seductive — but response variance is systematically understated, minority and tail positions get erased, stated willingness-to-pay is unstable under trivial re-phrasings of the prompt, and the errors are largest for exactly the populations least represented in training data. Decent mimics of the average; unreliable witnesses to the distribution.
⚠️
A synthetic panel cannot surprise you — and surprise is the product. Research earns its budget on the answers you didn't see coming; a model tuned to produce the plausible is structurally incapable of the implausible-but-true. If a study can only ever return answers you'd have guessed, it isn't research. It's autocomplete with a margin of error.
"New Coke won every sip test and lost the country."
The Coca-Cola Company · Project Kansas · 1985
the method200,000 blind taste tests, run competently. Sweeter won.
wrong questionA sip measures the liquid. The purchase was identity, ritual, ownership.
the result79 days of protest, hotlines melting, Classic hauled back in July.
the lessonMethod can be flawless while the question is wrong. Rigor doesn't rescue aim.
"Wrong answers, now instant and at scale."
LLM persona panels · 2026
the methodPrompt a thousand "consumers" and survey them in an afternoon, for pennies.
failure modeDistribution collapse: every persona drifts to the plausible middle; the weird tails vanish.
the echoYou're polling the training data — the cache again, cosplaying as a fresh sample.
the ruleFine for drafting questions. Inadmissible for demand, price, or anything a budget rides on.

8An insight pipeline that doesn't lie

Put the chapter together and you get a small set of operating rules — a pipeline, not a project.

Rule one: no decision on one source. Every stream in this chapter lies in a known direction — surveys flatter, search complains, reviews polarize, focus groups perform, synthetic panels average. Triangulation is the fix: a finding counts when streams with different biases agree on it. One source is an anecdote. Two is a hypothesis. Three, with unrelated failure modes, is a finding.

Rule two: the pipeline runs behavior → explanation → experiment. Behavioral data proposes ("trial users from channel X churn at twice the rate"). Qual explains ("the switch interviews say they were hiring us for a job we don't do"). An experiment confirms ("reposition the landing page for the real job; churn halves in the test cell"). Each stage covers the last one's blind spot: behavior can't tell you why, interviews can't tell you how many, and only the experiment tells you what happens if you act.

Rule three: write insights as falsifiable statements with an evidence grade attached. Not "users value simplicity" — that's Dichter's mistress, un-wrong-able and therefore worthless. Instead: "commuting trial users who hit feature X in week one retain 30% better — grade B: converging behavioral data plus interviews, no experiment yet." Grade A is a replicated experiment; B is converging behavioral streams; C is qual signal; D is opinion, however senior, and synthetic output lives at D. The grade travels with the claim, so a slide can't launder a hunch into a fact by formatting it nicely.

That's the whole discipline: sensors you don't fully trust, arranged so their lies cancel. Place the eight exhibits below where they belong on the map and see how your instincts compare to the chapter's grading.

Interactive · the triangulation board Click a chip, then a cell · place all 8 · then reveal
0 / 8 placed. Click a chip in the tray, then click a map cell.

One last thing the pipeline does, quietly: it tells you when to stop optimizing and start building. Sometimes every stream converges on the same awkward reading — the ladder from Chapter 1 is full, the incumbents own every rung that matters, and yet the switch interviews keep surfacing a job nobody serves. That's not a research finding to file. That's a door. The next chapter is about walking through it: when you can't win the existing category, you design a new one — and then you have to survive the gap in the middle of its adoption curve.

Check yourself

3 questions · instant feedback 0 / 3