Asked how probable something is, your fast system quietly answers a different question: how well does it match my stereotype? The fix isn't to stop seeing types — it's to start from the count.
Meet Tom W, a graduate student at your state university. A psychologist wrote this sketch of him years ago, during Tom's senior year of high school, based on personality tests of doubtful validity:
Now rank nine graduate fields by how probable it is that Tom is in each. When Kahneman and Tversky ran this, the ranking was nearly unanimous: computer science and engineering on top, humanities and education at the bottom. You just did it too — the sketch is a CS student, circa central casting.
Here's the problem. Enrollment numbers ran almost exactly the other way: humanities and education programs were huge, computer science was tiny. A random grad student was several times more likely to be in education than in CS. And the people ranking — including psychology grad students who could have recited those enrollment figures — ignored them completely.
Nobody was asked "how similar is Tom to your image of a CS student?" That was simply the question everyone answered. It's chapter 3's substitution move, running on rails: the hard question "how probable?" gets quietly swapped for the easy one "how similar to my stereotype?" — and the answer to the easy one comes back stamped as the answer to the hard one. Judging probability by resemblance has a name: the representativeness heuristic.
Suppose 3% of grad students are in computer science. Then a grad student you know nothing about is 3% likely to be one. That number — what you'd guess before hearing anything about this particular case — is the base rate. For the Bayesians in the room: it's the prior, and you already treat it with respect in code review. Here you throw it out the window.
The base rate should be your starting point. A flimsy, decade-old personality sketch should nudge it — corny puns are, let's grant, some evidence of CS membership — not replace it. But that's exactly what happens: the sketch doesn't adjust the 3%, it deletes it.
Why does the sketch win so completely? WYSIATI, from chapter 3: what you see is all there is. System 1 builds the best story from the information on the table. The sketch is vivid, specific, and sitting right in front of you. The base rate is a statistic you'd have to go look up — it never makes it to the table, so the story gets built without it. The machine isn't weighing sketch against statistic and choosing the sketch. The statistic was never in the room.
Quiet, loves books, wears glasses — librarian or farmer?
Librarian, obviously. The description is a librarian — the match is instant, vivid, and arrives pre-stamped as an answer. Similarity did all the work, felt like probability, and closed the case. Notice the question that never came up: how many male librarians are there, actually?
Male farmers outnumber male librarians in the US by roughly 20 : 1. Grant the stereotype everything it asks: say a bookish, quiet temperament is 4× as likely in a librarian as in a farmer. Run the odds: 1:20 × 4 = 4:20 — still five-to-one farmer. The bookish farmer wins on sheer headcount.
The transferable habit: start from the count, then let the sketch move you — not the other way around. A great personality match loses to a 20:1 base rate more often than not.
The most famous vignette in the history of judgment research. Read it, then answer honestly — the widget won't tell anyone.
Across every population Tversky and Kahneman tried — undergrads, statistically trained grad students, doctoral students in Stanford Business School's decision-science program — roughly 85 to 90% ranked "feminist bank teller" as more probable than "bank teller." Which is not a matter of opinion; it's a matter of set inclusion. Every feminist bank teller is a bank teller. Option B is a strict subset of option A. A conjunction cannot be more probable than either of its parts, for the same reason xs.filter(p) cannot be longer than xs — a filter only removes.
So what happened? Adding "active in the feminist movement" made the description a much better story about Linda — a much better match for the sketch — while making it strictly less probable. The story got better while the probability got worse, and the story won 9 times out of 10. That's plausibility impersonating probability, and it has a name: the conjunction fallacy.
Linda has been attacked for forty years, most persistently by Gerd Gigerenzer's group — and the attacks were not nonsense, so let's be honest about them. You just met the strongest one in the widget above: ask the question as counts ("of 100 women like Linda, how many are…") instead of asking which is "probable," and far fewer people commit the fallacy. Critics also noted that in everyday English "probable" can shade into "plausible" or "makes a good story" — so maybe subjects were answering a reasonable reading of an ambiguous word.
Judging by resemblance commits two systematic errors, and you've now seen the first: overweighting flimsy evidence that tells a good story. The Tom W sketch was explicitly stale and unreliable, and it steamrolled the enrollment numbers anyway. Quality of evidence never gets a vote; vividness does.
The second sin is insensitivity to sample size. Try this one:
Most people answer "about the same" — both hospitals see the same 50% coin, so why would they differ? But 15 births is a small sample, and small samples swing hard. Nine boys out of 15 is a perfectly ordinary Tuesday; 28 out of 45 is a genuinely rare day. Intuition treats the two hospitals as equally trustworthy summaries of the same process. They are not remotely.
Kahneman and Tversky called the underlying intuition "belief in the law of small numbers": the tacit assumption that a small sample is a miniature of its population, faithfully carrying all its properties. It isn't — small samples produce extreme results routinely, in both directions, meaning nothing. The uncomfortable punchline of their 1971 paper: professional research psychologists, people whose livelihood is statistical inference, chose sample sizes for their own experiments as if the law of small numbers were true. Trained intuition failed exactly like untrained intuition. The sin isn't ignorance; it's that resemblance-thinking doesn't have a slot for n.
The discipline is two moves, and neither requires notation:
That's Bayes' rule without the notation. Play the two moves against each other below — the dots are 100 concrete grad students, because as section 4 established, counts are the phrasing your System 2 actually shows up for.
For the record, the notation — once:
-- Bayes' rule, odds form: the whole discipline in one line posterior_odds = prior_odds × likelihood_ratio -- where LR = P(evidence | in category) / P(evidence | not) -- Tom W: prior 3:97, nerd sketch ~4× likelier for a CS student (3/97) × 4 ≈ 0.124 → 0.124 / 1.124 ≈ 11% -- (the box above rounds to whole students, so it says 12%)
One last honesty check. Stereotypes are often directionally right — CS really does have more sci-fi readers than education does; librarians really are more bookish than farmers on average. Representativeness isn't stupid; that's exactly why it's dangerous. The error isn't the direction, it's the magnitude: the sketch that deserves to move you from 3% to 11% instead moves you to "top of the ranking." The fix is not to pretend the sketch means nothing. The fix is arithmetic.