Anyone can run an A/B test — the tooling made it a checkbox. Running a test that tells you the truth, instead of what you hoped to hear, is a discipline. This chapter is the discipline.
Here's the whole chapter in one line: an A/B test is a machine for firing your opinions — and every shortcut you take quietly re-hires them. Everything below is the list of shortcuts, how each one lies, and what the honest version costs.
Every product meeting contains a moment where two smart people disagree about what customers will do, and the room resolves it by seniority. The A/B test exists to delete that moment. Split the audience at random, show each half a different version, count. The control group doesn't care who has the corner office — which is precisely its value. In the trade the phenomenon being deleted has a name: the HiPPO, the highest-paid person's opinion, and the entire experimentation movement is a machine for replacing it with a number.
None of this is new. Chapter 15 watched Claude Hopkins do it in 1923 with keyed coupons and regional split runs — "almost any question can be answered, cheaply, quickly and finally, by a test campaign." What a century added isn't the idea; it's the infrastructure. Feature flags, automatic stats, dashboards: the marginal cost of running a test fell to roughly zero. And here is the turn the century took that Hopkins didn't see coming — when testing became free, cheating became free too.
Hopkins' coupon test was honest by construction: one variable, one count, arithmetic a child could audit. A modern experimentation platform will happily let you check results hourly, swap the success metric mid-flight, and stop the moment the chart turns green. Each of those conveniences is a small machine for generating false conclusions. The modern failure mode is almost never no test. It's a test that looks rigorous and isn't — a ritual with a p-value on top.
So this chapter is not "why test" — Part V takes that as settled. It's the harder material: what an honest test is made of, the specific ways an honest-looking one lies, and what the shops that run tens of thousands of tests a year learned about how often anyone's ideas actually work. Fair warning: the base rate is humbling.
An honest experiment is a bet written down before the race. Five parts, all fixed before the first visitor arrives:
Programmer's version: a test spec written after reading the diff isn't a spec. You wouldn't trust a unit test authored by staring at the implementation until something passed; that's what an experiment becomes the moment the hypothesis, the metric, or the stopping rule is chosen after the data starts arriving.
There's a cheap pre-flight check that catches most dishonest tests before they run: ask what result would change the decision. If the feature ships regardless — because the exec sponsored it, because the redesign is already announced — then the test is decoration, and its cost is real: traffic, calendar time, and a false reputation for rigor. Cancel it and spend the sample somewhere a decision actually hangs on the answer.
The classical test you inherited from the textbook makes a specific promise: if there's no real difference, this procedure will fool you only 5% of the time. Every sin in this section works the same way — it quietly runs the procedure more than once while claiming the error rate of running it once. The 5% budget is per question. Ask many questions, spend many budgets.
Peeking. The dashboard updates hourly, so you look daily, and you stop the test the first morning the banner turns green. Mechanically: each look is a fresh chance for noise to wander across the significance line, and noise wanders — a test with no true effect drifts in and out of "significance" on its way to nowhere. Check a month-long test every day and your real false-positive rate isn't 5%; it's in the mid-20s. You built a smoke detector that samples the air once; you're pressing its test button forty times and reporting every beep as a fire.
Metric shopping. The primary metric came back flat, but engagement-on-mobile-for-new-users is up 12% — ship it? Twenty metrics at a 5% error rate means one false winner per test, by construction. The platform will always find you a green number; that's not insight, that's arithmetic. The honest version was decided in section 2: one OEC, named in advance. Secondary metrics generate the next hypothesis, never this test's verdict.
HARKing — hypothesizing after the results are known. The test "won" in a segment nobody mentioned beforehand ("it works for returning users on tablets!"), and a story is retro-fitted to the finding. This is drawing the target around the arrow. Segments are where noise goes to look like signal: slice 200 users twenty ways and several slices will "win" by luck alone. A segment result is a lead for the next pre-registered test — it is not a result.
Notice what all three sins have in common: no one is lying, no number is fabricated, and every individual step feels reasonable. The dishonesty is structural — it lives in when the questions were chosen, not in the answers. Which is why the fix is boring and absolute: decide first, then look. The widget below runs the experiment on you.
Here is the subtler failure, the one that poisons even tests run by the rules. Suppose the true effect of your change is a modest +2%, and your traffic is thin. A thin test is a noisy ruler: any single measurement lands somewhere in a wide band around the truth. Significance draws a line far out in that band and ships only what crosses it. Now follow the logic: if the line sits at +27% and the truth is +2%, the only way your test "wins" is by drawing a wildly lucky, wildly exaggerated measurement. The filter doesn't just select winners — it selects overstatements. Statisticians call it the significance filter; auction theorists, seeing the same shape, call it the winner's curse.
The consequence shows up two quarters later, in the ledger. Every shipped test claimed +8%, +15%, +30%; the sum of the claims says the business should have grown 40%; the business grew 4%. Nobody lied. The dashboard summed a set of estimates that were each individually filtered for luck — plus novelty effects (anything new gets clicked for a week, then the feed-trained brain re-tunes and the lift evaporates) and interactions between overlapping tests. Mature shops re-run their winners precisely to watch the shrinkage: the replication almost always lands closer to zero than the debut did.
The working defense is a slogan from the trade worth framing: Twyman's law — any figure that looks interesting or different is usually wrong. A surprising +30% from a small test is not a rocket; it's a bug report about your instrumentation, your randomization, or your luck. The bigger the surprise and the thinner the data, the harder you should shrink the estimate toward zero before you believe it — and the more valuable a boring, well-powered replication becomes.
The shops that industrialized testing — Microsoft, Booking.com, Airbnb, Netflix, running thousands of experiments a year for a decade — accidentally produced something more valuable than any single result: a base rate for human judgment about products. It is not flattering. Across tens of thousands of well-run tests, roughly a third of changes help, a third do nothing, and a third actively hurt. At Bing and Google, only 10–20% of ideas move their target metric. Booking.com's public number is blunter: about nine in ten experiments fail to beat control.
Two details make the number sting more. These are not random ideas — they're the shortlist: proposals that survived design review, resourcing, and a team's conviction, championed by people with domain expertise and skin in the game. And the experts can't pick the winners in advance: when experimentation teams ask staff to predict outcomes, accuracy hovers near coin-flip. Read that as the case for testing, not against ideas — if experts could predict results, you wouldn't need the tests. They can't, so you do.
The folklore cuts both ways. The famous Bing story — an ad-headline change parked in the backlog for months, worth ~$100M a year when it finally ran — is usually told as "small tweaks, huge wins." The honest reading is a pair of truths: big effects do occasionally hide in cheap changes (which is why you test broadly, ch15's spirit), and such outcomes are freak events against a base rate of flat (which is why you never budget for them).
The base rate also settles the small-site question. With 2,000 visitors a month, your test can only detect a near-doubling of conversion — and the base rate says near-doublings from tweaks essentially don't exist. So stop testing button shades on boutique traffic. Test big swings — a rewritten page, a different offer, a price — where a detectable effect is at least plausible, and take the button shade as a free choice, because statistically that's what it is. Match the boldness of the change to the resolution of the instrument; the widget below makes you practice.
Everything so far assumed the textbook shape: fix a sample size, wait, look once. Two legitimate relaxations exist, and both work by paying for what the sins tried to steal.
Sequential testing is peeking made legal. The insight is bookkeeping: if you want to look at the data repeatedly, budget for it — spend a sliver of your 5% error allowance at each interim look, with thresholds that start strict and relax as evidence accumulates. You buy the right to stop early (a genuinely large effect can be called in days) at the price of slightly larger samples when effects are small. Every serious platform now offers some flavor of this. The crime was never looking; it was looking without paying.
Bandits change the question. A multi-armed bandit shifts traffic toward whichever variant is currently winning — earn while you learn. For short-lived decisions where the answer stops mattering when the campaign ends (subject lines, headline rotations, promo banners), that's exactly right: you want maximum reward during the flight, not a publishable estimate after it. But note what you traded away: the losing arms get starved of traffic, so you end with a fuzzy read on how much better the winner was. Bandits optimize; experiments explain. Use bandits when you'll never need the number again; use tests when the number is the point.
The third complication doesn't relax the rules — it breaks them. Randomization assumes the arms don't touch: my seeing version B doesn't change what version A feels like to you. In marketplaces and social products that assumption fails structurally. Give half your riders a discount and they hire the same drivers the control group wanted — the treatment leaks into the control through the shared pool, flattering the test. Feeds, auctions, referral loops: same leak, different pipe. The fix is to randomize at the level where the interference stops — whole cities, whole markets — which hands you tiny sample sizes and a familiar shape: the geo experiment, star of the next chapter.
Be honest about the instrument's edges. Three things user-level testing structurally can't see:
When the unit test can't reach the behavior, you move up the pyramid: A/B test → geo experiment / holdout (ch29) → model (ch30). Same epistemology — compare against what would have happened otherwise — at coarser resolution and higher stakes. Part V is really one long escalation ladder, and you've just climbed the first rung.
Last thing, and it's the one that separates experimentation cultures from experimentation theaters: the notebook. Every test logged — hypothesis, design, result, decision — including the flat ones. Especially the flat ones. A failed test that's written down is a fence around a dead end, permanent institutional knowledge for the price you already paid; a failed test that vanishes gets re-run by the next team, and the one after. The shops with the humbling win rates aren't embarrassed by the two-thirds that didn't work. The two-thirds are the map — and the map is the asset the slot-machine crowd never accumulates.