You see a sliver of something and have to guess the whole — how long it'll run, how big it is, whether it'll happen again. The trick isn't more data; it's knowing what kind of thing you're looking at.
Chapters 3 through 5 were about handling what's already on your plate — sorting it, caching it, scheduling it. This chapter is about the stuff that hasn't happened yet. You get one glimpse of something and someone asks you to call the ending.
Here's the strange part: people are shockingly good at this. Tell someone a man is 39 and ask how long he'll live; tell them a movie has made $6 million and ask its total. They answer instantly, and they're roughly right. They're not doing statistics. They're doing something better, and it has a name.
A bag of coins. Some are fair; some are crooked — weighted heads, weighted tails, you don't know the mix. You reach in, pull one out, flip it once. Heads.
Now: what are the odds the next flip is heads too?
Your gut already answered. It said something like "a bit more than half — but not much more." Notice how specific that is. You didn't say 50%, because the coin did just land heads and that's weak evidence it leans that way. You didn't say 100%, because one flip is one flip and most coins in most bags are roughly fair. You landed in a narrow band without doing any arithmetic at all.
That instinct is the entire chapter. It has a formula, it has a name, and — this is the part worth sitting with — your gut was already computing it. What follows isn't a new skill. It's the manual for one you've been running unconsciously since you were three.
Here's the whole idea in one sentence: your guess after seeing evidence is what you believed before, bent by how surprising the evidence is. Not replaced by. Bent. The evidence gets a vote, not a veto, and how big its vote is depends on how weird it would be under each story you're entertaining.
The cleanest demonstration is medical, and it's the one that reliably breaks people's intuition. A disease affects 1 in 10,000 people. There's a test that's 99% accurate. You take it. It comes back positive. How worried should you be?
Most people say "99% worried." Let's not argue about it — let's just count. Line up 10,000 people and test every one of them.
So we have roughly 101 positive results in the room, and exactly one of them is real. If you're holding a positive slip, you're one of 101 people holding a positive slip, and your odds of being the unlucky one are about 1%. You are, overwhelmingly, fine.
Nothing was wrong with the test. The test did its job. What happened is that "1 in 10,000" was doing enormous work before the test showed up — and a 99%-accurate test simply isn't strong enough to overturn odds that lopsided. The healthy majority is so vast that even its tiny error rate produces more positives than the disease itself does. Drag the base rate in the widget below and watch the coral swamp the indigo.
Only now is the formula worth writing down, because now it's just bookkeeping for something you've already understood. This is Bayes's rule, published posthumously in 1763 on behalf of a Presbyterian minister named Thomas Bayes who apparently didn't think it was interesting enough to publish himself:
Read it left to right against the counting we just did. P(H) — the prior — is the 1-in-10,000. P(E|H) is the 99%: how likely a positive result is if you're sick. P(E) is all 101 positives, real and spurious. And P(H|E) — the posterior — is the 1%. The formula didn't tell you anything the dots didn't. It just does it faster, and without a room full of volunteers.
Back to the coin. You've flipped it three times now. Heads, heads, heads.
There are two ways to be wrong here, and both are popular. The naive optimist divides 3 by 3, announces "100% heads," and bets the house. The paralytic folds their arms and says we don't have enough data, come back at n=1000, refusing to produce any number at all. One of them is reckless and one of them is useless, and — this is the annoying bit — the paralytic usually thinks they're the rigorous one.
Pierre-Simon Laplace, working in the 1770s, gave the answer that splits the difference and turns out to be right. Count your successes and your attempts. Then add one to the top and two to the bottom:
Three heads out of three flips → 4/5 = 80%. Confident, not certain. Nine out of ten → 10/12 ≈ 83%, barely above the three-flip answer, because ten flips is still not much. Nine hundred out of a thousand → 901/1002 ≈ 89.9%, which is indistinguishable from the naive 90%. That's the elegance of it: the +1/+2 is a thumb on the scale that fades as evidence accumulates. It matters enormously at n=3 and disappears entirely by n=3000.
Where do the magic numbers come from? They're not magic. They're a prior — specifically, the flattest possible one, the belief that before you saw anything, every rate from 0 to 1 was equally plausible. Bake that in, turn the Bayesian crank, ask for the average, and out falls (s+1)/(n+2). It's Bayes's rule in a trench coat. That flat prior has a name too — Beta(1, 1) — and the code card later shows why updating it is a one-liner.
Suppose you know nothing. Not the bag, not the base rate, not the shape of anything. You're handed a single fact — this thing has been going for 20 years — and asked how much longer it'll last. Is there any principled answer at all?
There is, and it's almost insultingly simple: assume you showed up at a completely unremarkable moment. Not the beginning, not the end — somewhere ordinary. Which, on average, means the middle. And if you're standing at the middle of the thing's life, then it has about as much left as it's already had.
Total ≈ double the current age. A 20-year-old company gets another 20. A play in its 6th week runs 6 more. That's it — that's the rule.
In 1969, a physicist named J. Richard Gott stood at the Berlin Wall, which was eight years old at the time, and did exactly this arithmetic. He predicted it would stand somewhere between 2⅔ and 24 more years. The Wall came down in 1989 — twenty years later, comfortably inside his window, and rather more precisely than most of the professional Kremlinologists managed with their vastly larger budgets and their vastly worse priors.
The reasoning has a name: the Copernican principle, after the man who suggested that Earth is not, in fact, the centre of anything. Generalised: you are probably not special, and neither is the moment you're standing in.
Copernicus works when you know nothing. But you almost never know nothing — you know what kind of thing you're looking at. And it turns out that knowledge sorts the world into three families, each of which answers "how much longer?" in a completely different direction.
Human lifespan. Height. How long a cake takes to bake. These things cluster around a natural scale — there's a typical value and reality doesn't stray far from it. Nobody is 400 years old. Nobody is 12 feet tall. The cake will be done in half an hour, give or take.
Prediction here is additive: take the typical value, adjust a bit for what you've seen. A 39-year-old will live to about 80 — because that's the scale. An 80-year-old will not live to 160; they'll live a few more years. Every year that passes eats into a fixed budget. Elapsed time is bad news.
Money. City populations. Book sales. War casualties. How long a film stays in cinemas. These have no natural scale at all — no "typical" city size, no typical fortune. The distribution is a vast pile of tiny values with a thin tail stretching out to absurdity, and the tail is where all the interesting mass lives.
Prediction is multiplicative: your best guess for the total is the current value times some constant. A film that's grossed $6M is headed for about $14M. A film that's grossed $600M is headed for well over a billion. Same rule, same multiplier — and it never runs out, because there's no ceiling to run into. Elapsed time is good news. Every extra week is evidence you're dealing with a monster.
A politician's time in office. A slot machine's next jackpot. When a radioactive atom decays. These are memoryless — and that word is doing exactly what it says. The process does not know how long you've been waiting, does not care, and will not reward you for it.
Prediction is constant: whatever the expected remaining wait was when you arrived, that's what it still is. You've been at the bus stop 20 minutes? Expected wait: the same as minute zero. Your investment bought you precisely nothing. Elapsed time is no news. It's the only distribution that is genuinely indifferent to your suffering.
| World | Rule | What "it's been 20 minutes" means | Examples |
|---|---|---|---|
| Normal-ish | additive · predict near the typical value | "Nearly over." Elapsed time is bad news — you're eating a fixed budget. | lifespan, height, baking time, a commute |
| Power-law | multiplicative · predict a multiple of so-far | "Settle in." Elapsed time is good news — you're probably in the tail. | box office, wealth, city size, book sales |
| Erlang-ish | constant · predict the same wait as always | "No news." Elapsed time tells you nothing whatsoever. | political terms, jackpots, radioactive decay |
That table is the payload of the chapter, so here it is again without the table: the same evidence gives opposite advice depending on which world you're in. "It's been going 20 minutes" means nearly over, settle in, or no news — and nothing about the 20 minutes tells you which. Only knowing the shape does.
Drag the handle. All three panels see the identical evidence and disagree completely.
So back to the opening puzzle: why are ordinary people so good at this? Nobody at a bus stop is integrating a posterior. And yet they get it right.
Tom Griffiths and Josh Tenenbaum ran the experiment properly. They gave people single data points across wildly different domains — a man's age, a movie's gross so far, the length of a poem, how long a pharaoh had reigned, how long a cake had been in the oven — and asked for the total. One number in, one number out, no context, no training.
The results are the good part. People didn't apply one rule everywhere. Asked about lifespans they predicted additively. Asked about box office they predicted multiplicatively. Asked about pharaohs — a domain almost none of them knew anything about — they produced the memoryless answer, which happens to be correct, because Egyptian reigns really did follow an Erlang distribution. Their guesses tracked the true distribution for each domain, one domain at a time.
Which tells you what the superpower actually is. It isn't inference — the inference is a division. It's that decades of living in the world have quietly installed, in each of us, a startlingly well-calibrated library of priors: a sense of what shape each corner of reality is. You don't need much data when you already know what you're looking at. "Small data" is just data with a good prior attached.
Here's the sting in the tail. Bayes's rule is a machine for combining evidence with priors, and it is scrupulously, mechanically honest. Which means that if your priors are garbage, it will carry the garbage through with perfect fidelity and hand you back a confident, well-reasoned, completely wrong belief.
And where do your priors come from? Not from a textbook. From counting. From whatever you happened to see. Your sense of "how often does X happen" is a tally of the Xs that crossed your field of view — which is fine, as long as your field of view is a fair sample of the world.
It isn't. If your window onto reality is a feed tuned to maximise engagement, you are not sampling the world; you are sampling the world's most alarming 0.001%, hand-picked and served warm. Plane crashes, not plane landings. Violent crime, not the Tuesday where nothing happened. The machinery in your head does exactly what it's supposed to with that input, and produces a person who is calibrated for a world that doesn't exist.
The advice that falls out of this is deeply unglamorous, which is usually a sign it's correct: your predictions are downstream of your inputs, so choose your inputs. Not "think harder" — thinking harder just applies Bayes more precisely to the same bad prior. The leverage isn't in the reasoning. It's upstream of it.
The remarkable thing about coding this up is how little there is. laplace is one division. And the general Bayesian update, for a coin, is just adding what you saw to what you believed — a Beta(a, b) goes to Beta(a + s, b + f) and that's the entire update. Which means believing things over a stream of evidence is a foldLeft. You've written this function a hundred times without knowing it was epistemology.
believe. It's foldLeft with a prior as the seed and evidence as the stream — the shape you already reach for without thinking. That's not an analogy. A Bayesian agent is a fold over its experience, carrying a belief as the accumulator. The prior is the seed value, and if you seed it wrong, no amount of folding saves you.This chapter argued that good priors let you predict from almost nothing. The next one is the flip side, and it's worse than it sounds: what happens when you take the data too seriously. There's a point where more thinking, more factors and more evidence stop helping and start actively making you worse — and the fix is to deliberately think less. It is not a comfortable chapter for anyone who likes being thorough.