🧠 Thinking, Fast and Slow · ch.7 · regression to the mean
🧠 Chapter 7 · Part II — Heuristics & Biases

The jinx is
just math

Every extreme result is part talent, part dice — and the dice get rolled again next time. This chapter is about the drift back toward normal that follows every peak and every disaster, and why your mind insists on inventing a cause for it.

1The flight instructors

The most important eureka of Kahneman's career happened in a room full of Israeli Air Force flight instructors — and it began with one of them telling him he was full of it.

Kahneman was lecturing on the psychology of training: reward for improvement works better than punishment for mistakes. A seasoned instructor stood up. In his experience it was exactly backwards. "I've praised cadets warmly for a beautifully executed maneuver — and the next time, they almost always do worse. I've screamed at cadets for a terrible one — and the next time, they almost always do better. So don't tell me reward works and punishment doesn't."

Here's the uncomfortable part: the instructor's observations were correct. Landings after praise really were worse, on average. Landings after screaming really were better. He was wrong only about the cause — and the true cause involved no psychology at all.

An unusually good landing isn't pure skill. It's skill plus a lucky draw — calm air, sharp focus, a well-timed flare. Next landing, the luck is drawn again, and it's usually less generous. So the landing after a great one tends to be worse — praised or not. An awful landing is partly bad luck, so the next one tends to be better — screamed at or not. If you're an engineer, you already know this pattern: rerun the flakiest CI run in the history of your repo, and it improves. You didn't fix anything. The flake got resampled.

The instructors had spent careers attaching a feedback policy to a statistical artifact. Name arrives last, as promised: regression to the mean.

💬
Kahneman's summary of what the Air Force taught him: "Because we tend to reward others when they do well and punish them when they do badly, and because there is regression to the mean, it is part of the human condition that we are statistically punished for rewarding others and rewarded for punishing them."
Interactive · the screaming instructor one squadron · 20 landings · pick a feedback policy

2Performance = skill + luck

One equation explains everything in this chapter: performance = skill + luck. The two terms have different lifetimes. Skill persists between observations; luck is drawn fresh every single time. In Scala terms, skill is a val — evaluated once, stable across calls. Luck is a def — re-evaluated at every call site. The whole chapter is about what happens when you forget the parentheses on luck().

Now think about what an extreme observation means. To post the best score in the field, it's not enough to be the most skilled — you also need the luck term to break your way on the same day. Extreme = skill and luck lined up. Next observation, the skill is still there, but the luck reverts to boring. So the day-1 leader usually posts a worse day-2 score — still good, just less spectacular — and the day-1 straggler usually improves. Both drift toward the field average, from opposite directions, for the same reason.

Kahneman's example is a golf tournament. A 66 on day 1 says two things: this golfer is better than average, and today went their way. Your best prediction for their day 2 isn't 66 — it's something between 66 and the field mean. Same logic, mirrored, for the golfer who shot 77.

One more reframe, and it's the deepest one. Take two noisy measurements of the same underlying thing — day-1 and day-2 scores, parent and child heights, this quarter and next quarter. Their correlation is less than 1, because noise doesn't correlate with itself. And "correlation < 1" and "extremes regress" are not two facts — they're the same fact read in two directions. Whenever two measures correlate imperfectly, extremes on one will be less extreme on the other. No mechanism required. That's the whole theorem.

Interactive · two days of golf 80 simulated players · score = skill + fresh luck each day
day-1 leaders day-1 stragglers everyone else golf: lower score = better

3The jinx dossier

Once you have the equation, a whole folklore of curses dissolves into one boring sentence. The case files:

  • The Sports Illustrated cover jinx. Athletes who make the cover promptly decline. Of course they do — you only make the cover after an extreme stretch, extreme means the luck term was maxed out, and luck doesn't repeat on schedule. The cover doesn't curse anyone; it selects for people about to regress.
  • The sophomore slump. Rookie of the year, disappointing second season. Same file: "rookie of the year" is an award for an extreme, and extremes are partly luck. The slump is mostly the debut's fault, not the sophomore's.
  • Manager of the year. Their team underperforms next season, and the business press writes think-pieces about complacency and lost hunger. The selection did the work: you don't win the award without an outlier year.
  • The hot fund. Last year's top-decile fund, this year's mediocrity. Hold that one — chapter 8 finishes it off properly, with data that should end a whole industry and somehow doesn't.

Notice what every entry has in common: a selection on an extreme, followed by surprise at the reversion, followed by a manufactured cause — jinx, complacency, pressure, hunger. Nobody ever files the correct report: "subject was selected because the dice came up sixes; dice have since returned to normal operation."

4Why you can't see it

If regression is this simple, why did humanity need until 1886 — Francis Galton, measuring parents and their children — to notice it? And why, a century and a half later, does your sprint retro still explain every metric wobble with a story?

Because of the machine you met in chapter 2. System 1 is a story-writer: give it an event and it will produce a cause, instantly, without being asked. But regression has no cause. It isn't a force. Nothing pulls the outlier back; no mechanism "acts on" the sophomore. Regression is just what "noisy" means, viewed across two samples. A fact with no causal story is a fact System 1 cannot represent — so it papers over the gap with screaming instructors, jinxes, and lost motivation. The math is invisible precisely because it's not a story.

And it took real machinery to see. Galton needed the toolkit of two centuries of post-Newton mathematics — and several years of his own confusion — before he understood what his height data was telling him. Your intuition, running on zero centuries of mathematics and a deadline, is not going to spot it between stand-up and lunch.

The trap has a standard industrial form, and you've read it in a hundred press releases: "depressed children who drank the energy drink for three months showed significant improvement." They did! Depressed children are, by construction, a group measured at an extreme. Measure them again later and they improve on average — energy drink, kale smoothie, or nothing at all. Any group selected for an extreme score improves on remeasurement. That is the entire reason treatment effects require control groups: the control arm exists to measure how much improvement regression hands you for free, so you can subtract it.

⚠️
Field guide: any claim of the form "we took the worst-performing X, applied Y, and X improved" is regression until proven otherwise. Worst-performing schools that got the grant, slowest services that got the task force, unhappiest customers who got the outreach call. Ask what the untouched extremes did.

5Green cabs, Blue cabs

Regression isn't the only statistical fact System 1 refuses to see. Here's Kahneman's favorite demonstration of the general disease — statistics without a story bounce right off. A cab was involved in a hit-and-run at night. Two facts:

  • 85% of the city's cabs are Green, 15% are Blue.
  • A witness identified the cab as Blue — and under similar night conditions, the witness gets the color right 80% of the time.

How likely is it the cab was actually Blue? Nearly everyone says something close to 80% — the witness's reliability, full stop. The 85% base rate might as well not exist. But run the count: the city is so Green that even a good witness's "Blue" calls are mostly mistaken Greens. The honest answer is 41% — the cab is more likely Green than Blue even after a credible witness said Blue. Don't take my word for it; count the cabs below.

Then comes the twist Kahneman loved. Reword one sentence: instead of "85% of cabs are Green," say "Green cab drivers cause 85% of the accidents." Same number, same arithmetic — and suddenly people use the base rate. Why? Because now it's a story: Green drivers are reckless. A statistical base rate — a fact about the population — gets ignored. A causal base rate — a fact that implies a trait — gets absorbed into System 1's narrative and duly applied. Stereotypes, which you met in chapter 6, are exactly this: causal base rates about groups. That's why they stick when statistics slide off.

Interactive · the cab problem, counted 100 cabs · step through the witness's testimony
step 0 / 3
2026 check Base-rate neglect replicates, but with an asterisk Gerd Gigerenzer spent a career sharpening: people ignore base rates stated as probabilities ("the witness is 80% reliable") far more than base rates stated as natural frequencies ("of 100 cabs, the witness calls 29 'Blue' and only 12 really are"). Formats that let you count concrete cases reach intuition; abstract percentages don't. That's not a footnote to the widget above — it's the reason the widget is a pile of countable dots instead of a formula.

6Living with regression

The operational rule: expect regression anywhere you selected on an extreme. Not "sometimes," not "in sports" — anywhere the thing that got your attention got it by being an outlier on a noisy measurement. A running list from the places you actually work:

  • The interview superstar. An interview is one noisy sample, and you hired the max of the batch. The max of a noisy batch is the most regression-prone number in statistics. Expect a merely-good employee, and stop treating the gap as a hiring failure or a motivation mystery.
  • The A/B test winner, re-run. You picked the variant because it topped the leaderboard, so its measured lift is skill plus favorable noise. Re-run it and the lift shrinks — the "winner's curse" of metrics. Budget for shrinkage before you promise the roadmap that +12%.
  • The breakthrough quarter. Whatever the team does next, plan for numbers closer to trend — and write that in the forecast before the merely-normal quarter arrives, so nobody has to invent a villain for it.

And the single portable habit: before crediting any cause for a change back toward normal — a fix, a pep talk, a reorg, a new vitamin — ask one question first. "How noisy is this measurement?" If the answer is "noisy" and the starting point was extreme, regression has already explained most of the change, and the burden of proof is on the story.

Your best engineer had a stellar quarter; this one is merely fine. Something changed. They've lost motivation since the promotion — or they're interviewing, or the new manager is smothering them. The story-writer needs a villain, and it will have one by end of stand-up. Find the cause, fix the person.

A stellar quarter is skill + a good luck draw — the right tickets, a cooperative dependency, a bug that happened to be shallow. The skill renews next quarter; the luck draw doesn't. A merely-fine follow-up isn't a decline from baseline; it usually is the baseline, seen after one lucky sample.

The transferable habit: before hunting causes for a dip from a peak, ask what the measurement's noise band looks like. Most "declines from a peak" are the peak's fault, not the person's.