Marketing — the Atlas · ch.28 · experimentation
📣 Chapter 28 · Part V · Growth & measurement

A test you can't lie to

Anyone can run an A/B test — the tooling made it a checkbox. Running a test that tells you the truth, instead of what you hoped to hear, is a discipline. This chapter is the discipline.

Here's the whole chapter in one line: an A/B test is a machine for firing your opinions — and every shortcut you take quietly re-hires them. Everything below is the list of shortcuts, how each one lies, and what the honest version costs.

1Testing beats arguing

Every product meeting contains a moment where two smart people disagree about what customers will do, and the room resolves it by seniority. The A/B test exists to delete that moment. Split the audience at random, show each half a different version, count. The control group doesn't care who has the corner office — which is precisely its value. In the trade the phenomenon being deleted has a name: the HiPPO, the highest-paid person's opinion, and the entire experimentation movement is a machine for replacing it with a number.

None of this is new. Chapter 15 watched Claude Hopkins do it in 1923 with keyed coupons and regional split runs — "almost any question can be answered, cheaply, quickly and finally, by a test campaign." What a century added isn't the idea; it's the infrastructure. Feature flags, automatic stats, dashboards: the marginal cost of running a test fell to roughly zero. And here is the turn the century took that Hopkins didn't see coming — when testing became free, cheating became free too.

Hopkins' coupon test was honest by construction: one variable, one count, arithmetic a child could audit. A modern experimentation platform will happily let you check results hourly, swap the success metric mid-flight, and stop the moment the chart turns green. Each of those conveniences is a small machine for generating false conclusions. The modern failure mode is almost never no test. It's a test that looks rigorous and isn't — a ritual with a p-value on top.

So this chapter is not "why test" — Part V takes that as settled. It's the harder material: what an honest test is made of, the specific ways an honest-looking one lies, and what the shops that run tens of thousands of tests a year learned about how often anyone's ideas actually work. Fair warning: the base rate is humbling.

One variable, one count, no way to cheat.
Claude Hopkins · Scientific Advertising · 1923
the methodTwo versions of an ad, two coupon codes, two piles of mail. The bigger pile wins. Chapter 15 called the coupon a conversion pixel made of paper.
the honestyIntegrity was built into the design: one look, at the end, at arithmetic anyone could audit. There was nothing to peek at and no metric to shop for.
the limitWeeks per answer, one question at a time, and only questions a coupon could carry.
the lessonThe test was slow and the honesty was free. The next hundred years reversed both.
A thousand tests, and a leaderboard watching.
The modern experimentation stack · flags, stats engines, dashboards
the machineFeature flags split traffic, the stats compute themselves, results stream to a dashboard in real time. Hopkins' dream, fully automated.
the winThe marginal cost of a test is near zero — so mature shops run thousands a year, on everything from ranking models to button copy.
the catchThe marginal cost of a bad test is also near zero. Peeking is the default view. Metric shopping is a dropdown. Stopping on green is one click.
the lessonThe arithmetic got automated; the incentives didn't. Honesty moved from the math into the process — which is why this chapter is about process.

2Anatomy of an honest test

An honest experiment is a bet written down before the race. Five parts, all fixed before the first visitor arrives:

  • A hypothesis, stated in advance. "Removing the second checkout step will raise completion, because the field-level analytics show 30% of drop-off happens there." Not "let's try stuff and see what the data says" — the data will say something either way; that's what noise does.
  • One primary metric. The experimentation literature calls it the OEC — the overall evaluation criterion. One number, agreed in advance, that decides the test. Everything else is commentary.
  • A pre-committed sample size. Compute how many visitors you need to detect the effect you plausibly expect, and run until you have them. The sample size is the test's contract with itself — the whole next section is about what happens when you break it.
  • Guardrail metrics. The numbers that aren't allowed to get worse while the primary improves: revenue per visitor, page speed, unsubscribes, support tickets. A "win" that quietly trades against a guardrail is a loss with good PR.
  • A randomization unit that matches the decision. Users, sessions, markets — pick the unit the treatment actually touches, and keep each unit in one arm for the duration.

Programmer's version: a test spec written after reading the diff isn't a spec. You wouldn't trust a unit test authored by staring at the implementation until something passed; that's what an experiment becomes the moment the hypothesis, the metric, or the stopping rule is chosen after the data starts arriving.

There's a cheap pre-flight check that catches most dishonest tests before they run: ask what result would change the decision. If the feature ships regardless — because the exec sponsored it, because the redesign is already announced — then the test is decoration, and its cost is real: traffic, calendar time, and a false reputation for rigor. Cancel it and spend the sample somewhere a decision actually hangs on the answer.

💡
Write four things down before the data arrives: the hypothesis, the primary metric, the sample size, and the result that kills the idea. That last line is the one that stings — and it's the one that makes the other three mean something. A test you can't lose isn't a test.

3The sins: peeking, shopping, HARKing

The classical test you inherited from the textbook makes a specific promise: if there's no real difference, this procedure will fool you only 5% of the time. Every sin in this section works the same way — it quietly runs the procedure more than once while claiming the error rate of running it once. The 5% budget is per question. Ask many questions, spend many budgets.

Peeking. The dashboard updates hourly, so you look daily, and you stop the test the first morning the banner turns green. Mechanically: each look is a fresh chance for noise to wander across the significance line, and noise wanders — a test with no true effect drifts in and out of "significance" on its way to nowhere. Check a month-long test every day and your real false-positive rate isn't 5%; it's in the mid-20s. You built a smoke detector that samples the air once; you're pressing its test button forty times and reporting every beep as a fire.

Metric shopping. The primary metric came back flat, but engagement-on-mobile-for-new-users is up 12% — ship it? Twenty metrics at a 5% error rate means one false winner per test, by construction. The platform will always find you a green number; that's not insight, that's arithmetic. The honest version was decided in section 2: one OEC, named in advance. Secondary metrics generate the next hypothesis, never this test's verdict.

HARKing — hypothesizing after the results are known. The test "won" in a segment nobody mentioned beforehand ("it works for returning users on tablets!"), and a story is retro-fitted to the finding. This is drawing the target around the arrow. Segments are where noise goes to look like signal: slice 200 users twenty ways and several slices will "win" by luck alone. A segment result is a lead for the next pre-registered test — it is not a result.

Notice what all three sins have in common: no one is lying, no number is fabricated, and every individual step feels reasonable. The dishonesty is structural — it lives in when the questions were chosen, not in the answers. Which is why the fix is boring and absolute: decide first, then look. The widget below runs the experiment on you.

Interactive · the peeking machine 200 A/A tests · both arms identical · every "win" is a false alarm
same 200 experiments · two stopping rules
Both arms serve the identical page — there is nothing to find. Run the batch and watch how many "discoveries" each judging rule produces from pure noise.

4The winner's curse

Here is the subtler failure, the one that poisons even tests run by the rules. Suppose the true effect of your change is a modest +2%, and your traffic is thin. A thin test is a noisy ruler: any single measurement lands somewhere in a wide band around the truth. Significance draws a line far out in that band and ships only what crosses it. Now follow the logic: if the line sits at +27% and the truth is +2%, the only way your test "wins" is by drawing a wildly lucky, wildly exaggerated measurement. The filter doesn't just select winners — it selects overstatements. Statisticians call it the significance filter; auction theorists, seeing the same shape, call it the winner's curse.

The consequence shows up two quarters later, in the ledger. Every shipped test claimed +8%, +15%, +30%; the sum of the claims says the business should have grown 40%; the business grew 4%. Nobody lied. The dashboard summed a set of estimates that were each individually filtered for luck — plus novelty effects (anything new gets clicked for a week, then the feed-trained brain re-tunes and the lift evaporates) and interactions between overlapping tests. Mature shops re-run their winners precisely to watch the shrinkage: the replication almost always lands closer to zero than the debut did.

The working defense is a slogan from the trade worth framing: Twyman's law — any figure that looks interesting or different is usually wrong. A surprising +30% from a small test is not a rocket; it's a bug report about your instrumentation, your randomization, or your luck. The bigger the surprise and the thinner the data, the harder you should shrink the estimate toward zero before you believe it — and the more valuable a boring, well-powered replication becomes.

⚠️
The dashboard that only goes up. Sum of shipped-win estimates: +40% for the year. Actual ledger: +4%. No fraud required — just the significance filter (only exaggerated draws get shipped), novelty decay, and twenty overlapping tests each claiming the same conversion. If your experiment program's claimed wins never reconcile against the ledger, the program isn't measuring the business; it's measuring its own optimism.
Interactive · the winner's curse what shipped "wins" report vs what's true
+2%
2k
Pick a true effect and a sample size. The histogram shows what the tests that reached significance reported — the shipped wins — against the truth.

5The base rate of good ideas

The shops that industrialized testing — Microsoft, Booking.com, Airbnb, Netflix, running thousands of experiments a year for a decade — accidentally produced something more valuable than any single result: a base rate for human judgment about products. It is not flattering. Across tens of thousands of well-run tests, roughly a third of changes help, a third do nothing, and a third actively hurt. At Bing and Google, only 10–20% of ideas move their target metric. Booking.com's public number is blunter: about nine in ten experiments fail to beat control.

Two details make the number sting more. These are not random ideas — they're the shortlist: proposals that survived design review, resourcing, and a team's conviction, championed by people with domain expertise and skin in the game. And the experts can't pick the winners in advance: when experimentation teams ask staff to predict outcomes, accuracy hovers near coin-flip. Read that as the case for testing, not against ideas — if experts could predict results, you wouldn't need the tests. They can't, so you do.

The folklore cuts both ways. The famous Bing story — an ad-headline change parked in the backlog for months, worth ~$100M a year when it finally ran — is usually told as "small tweaks, huge wins." The honest reading is a pair of truths: big effects do occasionally hide in cheap changes (which is why you test broadly, ch15's spirit), and such outcomes are freak events against a base rate of flat (which is why you never budget for them).

The base rate also settles the small-site question. With 2,000 visitors a month, your test can only detect a near-doubling of conversion — and the base rate says near-doublings from tweaks essentially don't exist. So stop testing button shades on boutique traffic. Test big swings — a rewritten page, a different offer, a price — where a detectable effect is at least plausible, and take the button shade as a free choice, because statistically that's what it is. Match the boldness of the change to the resolution of the instrument; the widget below makes you practice.

📈
Evidence check. The canonical industry figures, all directional but remarkably consistent: Microsoft's experimentation group reports roughly ⅓ of changes help, ⅓ are flat, ⅓ hurt; Bing/Google-scale teams see only 10–20% of ideas win; Booking.com has said publicly that ~9 of 10 experiments fail. On the sins: simulation of a month-long test checked daily pushes the nominal 5% false-positive rate to ~25–30%. And replications of significant wins consistently show shrinkage toward zero — the winner's curse, measured. If your program's win rate is 70%, the polite hypothesis is that your tests are underpowered and your winners exaggerated; the impolite one is that somebody's peeking.
Interactive · the test triager six proposals · triage each before the sprint starts
verdicts test it now bigger swing first can't A/B — escalate 0 / 6 triaged
For each proposal: is the instrument sharp enough for the effect this change could plausibly have? If yes, test. If no, swing bigger. If no control group can exist, escalate to Chapter 29's tools.

6Beyond the fixed horizon

Everything so far assumed the textbook shape: fix a sample size, wait, look once. Two legitimate relaxations exist, and both work by paying for what the sins tried to steal.

Sequential testing is peeking made legal. The insight is bookkeeping: if you want to look at the data repeatedly, budget for it — spend a sliver of your 5% error allowance at each interim look, with thresholds that start strict and relax as evidence accumulates. You buy the right to stop early (a genuinely large effect can be called in days) at the price of slightly larger samples when effects are small. Every serious platform now offers some flavor of this. The crime was never looking; it was looking without paying.

Bandits change the question. A multi-armed bandit shifts traffic toward whichever variant is currently winning — earn while you learn. For short-lived decisions where the answer stops mattering when the campaign ends (subject lines, headline rotations, promo banners), that's exactly right: you want maximum reward during the flight, not a publishable estimate after it. But note what you traded away: the losing arms get starved of traffic, so you end with a fuzzy read on how much better the winner was. Bandits optimize; experiments explain. Use bandits when you'll never need the number again; use tests when the number is the point.

The third complication doesn't relax the rules — it breaks them. Randomization assumes the arms don't touch: my seeing version B doesn't change what version A feels like to you. In marketplaces and social products that assumption fails structurally. Give half your riders a discount and they hire the same drivers the control group wanted — the treatment leaks into the control through the shared pool, flattering the test. Feeds, auctions, referral loops: same leak, different pipe. The fix is to randomize at the level where the interference stops — whole cities, whole markets — which hands you tiny sample sizes and a familiar shape: the geo experiment, star of the next chapter.

7What you can't A/B — and the notebook

Be honest about the instrument's edges. Three things user-level testing structurally can't see:

  • Brand. Chapter 19's long clock: brand advertising works over quarters, on people who aren't buying yet, through channels everyone sees at once. There is no user-level control group for a billboard, and a 28-day test window can't price an asset that pays out over three years.
  • Tiny samples. A B2B motion closing 30 deals a quarter cannot A/B its pricing page — fifteen conversions per arm answers nothing this year. Decide with research (ch10), price with judgment (ch09), and monitor guardrails instead of pretending to test.
  • Long horizons. Retention at month 12, LTV effects of onboarding, the cost of a trust-burning dark pattern (ch13's ledger) — the test ends years before the truth arrives. Cohort tracking and holdouts, not dashboards.

When the unit test can't reach the behavior, you move up the pyramid: A/B test → geo experiment / holdout (ch29) → model (ch30). Same epistemology — compare against what would have happened otherwise — at coarser resolution and higher stakes. Part V is really one long escalation ladder, and you've just climbed the first rung.

Last thing, and it's the one that separates experimentation cultures from experimentation theaters: the notebook. Every test logged — hypothesis, design, result, decision — including the flat ones. Especially the flat ones. A failed test that's written down is a fence around a dead end, permanent institutional knowledge for the price you already paid; a failed test that vanishes gets re-run by the next team, and the one after. The shops with the humbling win rates aren't embarrassed by the two-thirds that didn't work. The two-thirds are the map — and the map is the asset the slot-machine crowd never accumulates.

The ugliest subject line raised the most money.
Obama for America digital team · 2012 cycle
the moveDraft a dozen-plus subject lines and bodies, test them on a slice of the list, send the winner to tens of millions — every send, all cycle.
the winner"Hey" and other homely, casual lines beat the polished copy, over and over. The staff's aesthetic favorites lost routinely.
the surpriseVeteran staffers guessed which lines would win — and were reliably wrong. Taste generated the candidates; only the list could rank them.
the lessonTesting attributed on the order of $200M in donations. Taste is a hypothesis generator, not a judge — the base rate from section 5, wearing a campaign T-shirt.
Two hundred variants before lunch.
Generative creative × experimentation platforms · 2026
the moveAI drafts hundreds of headlines, hooks, and layouts on demand; the platform can test them all. Generation is no longer the bottleneck anywhere.
the math problemTwenty variants at 95% confidence expects one false winner by construction. Two hundred expects ten — a leaderboard of lucky noise, refreshed daily.
the fixMultiple-comparison corrections, bandits for the short-lived stuff, and one iron rule: a variant that "won" once is a lottery ticket until it replicates on fresh traffic.
the lessonGeneration got cheap; confirmation didn't. The scarce resource in 2026 is the same one Hopkins rationed in 1923 — honest sample size.

Check yourself

3 questions · instant feedback 0 / 3