Every extreme result is part talent, part dice — and the dice get rolled again next time. This chapter is about the drift back toward normal that follows every peak and every disaster, and why your mind insists on inventing a cause for it.
The most important eureka of Kahneman's career happened in a room full of Israeli Air Force flight instructors — and it began with one of them telling him he was full of it.
Kahneman was lecturing on the psychology of training: reward for improvement works better than punishment for mistakes. A seasoned instructor stood up. In his experience it was exactly backwards. "I've praised cadets warmly for a beautifully executed maneuver — and the next time, they almost always do worse. I've screamed at cadets for a terrible one — and the next time, they almost always do better. So don't tell me reward works and punishment doesn't."
Here's the uncomfortable part: the instructor's observations were correct. Landings after praise really were worse, on average. Landings after screaming really were better. He was wrong only about the cause — and the true cause involved no psychology at all.
An unusually good landing isn't pure skill. It's skill plus a lucky draw — calm air, sharp focus, a well-timed flare. Next landing, the luck is drawn again, and it's usually less generous. So the landing after a great one tends to be worse — praised or not. An awful landing is partly bad luck, so the next one tends to be better — screamed at or not. If you're an engineer, you already know this pattern: rerun the flakiest CI run in the history of your repo, and it improves. You didn't fix anything. The flake got resampled.
The instructors had spent careers attaching a feedback policy to a statistical artifact. Name arrives last, as promised: regression to the mean.
One equation explains everything in this chapter: performance = skill + luck. The two terms have different lifetimes. Skill persists between observations; luck is drawn fresh every single time. In Scala terms, skill is a val — evaluated once, stable across calls. Luck is a def — re-evaluated at every call site. The whole chapter is about what happens when you forget the parentheses on luck().
Now think about what an extreme observation means. To post the best score in the field, it's not enough to be the most skilled — you also need the luck term to break your way on the same day. Extreme = skill and luck lined up. Next observation, the skill is still there, but the luck reverts to boring. So the day-1 leader usually posts a worse day-2 score — still good, just less spectacular — and the day-1 straggler usually improves. Both drift toward the field average, from opposite directions, for the same reason.
Kahneman's example is a golf tournament. A 66 on day 1 says two things: this golfer is better than average, and today went their way. Your best prediction for their day 2 isn't 66 — it's something between 66 and the field mean. Same logic, mirrored, for the golfer who shot 77.
One more reframe, and it's the deepest one. Take two noisy measurements of the same underlying thing — day-1 and day-2 scores, parent and child heights, this quarter and next quarter. Their correlation is less than 1, because noise doesn't correlate with itself. And "correlation < 1" and "extremes regress" are not two facts — they're the same fact read in two directions. Whenever two measures correlate imperfectly, extremes on one will be less extreme on the other. No mechanism required. That's the whole theorem.
Once you have the equation, a whole folklore of curses dissolves into one boring sentence. The case files:
Notice what every entry has in common: a selection on an extreme, followed by surprise at the reversion, followed by a manufactured cause — jinx, complacency, pressure, hunger. Nobody ever files the correct report: "subject was selected because the dice came up sixes; dice have since returned to normal operation."
If regression is this simple, why did humanity need until 1886 — Francis Galton, measuring parents and their children — to notice it? And why, a century and a half later, does your sprint retro still explain every metric wobble with a story?
Because of the machine you met in chapter 2. System 1 is a story-writer: give it an event and it will produce a cause, instantly, without being asked. But regression has no cause. It isn't a force. Nothing pulls the outlier back; no mechanism "acts on" the sophomore. Regression is just what "noisy" means, viewed across two samples. A fact with no causal story is a fact System 1 cannot represent — so it papers over the gap with screaming instructors, jinxes, and lost motivation. The math is invisible precisely because it's not a story.
And it took real machinery to see. Galton needed the toolkit of two centuries of post-Newton mathematics — and several years of his own confusion — before he understood what his height data was telling him. Your intuition, running on zero centuries of mathematics and a deadline, is not going to spot it between stand-up and lunch.
The trap has a standard industrial form, and you've read it in a hundred press releases: "depressed children who drank the energy drink for three months showed significant improvement." They did! Depressed children are, by construction, a group measured at an extreme. Measure them again later and they improve on average — energy drink, kale smoothie, or nothing at all. Any group selected for an extreme score improves on remeasurement. That is the entire reason treatment effects require control groups: the control arm exists to measure how much improvement regression hands you for free, so you can subtract it.
Regression isn't the only statistical fact System 1 refuses to see. Here's Kahneman's favorite demonstration of the general disease — statistics without a story bounce right off. A cab was involved in a hit-and-run at night. Two facts:
How likely is it the cab was actually Blue? Nearly everyone says something close to 80% — the witness's reliability, full stop. The 85% base rate might as well not exist. But run the count: the city is so Green that even a good witness's "Blue" calls are mostly mistaken Greens. The honest answer is 41% — the cab is more likely Green than Blue even after a credible witness said Blue. Don't take my word for it; count the cabs below.
Then comes the twist Kahneman loved. Reword one sentence: instead of "85% of cabs are Green," say "Green cab drivers cause 85% of the accidents." Same number, same arithmetic — and suddenly people use the base rate. Why? Because now it's a story: Green drivers are reckless. A statistical base rate — a fact about the population — gets ignored. A causal base rate — a fact that implies a trait — gets absorbed into System 1's narrative and duly applied. Stereotypes, which you met in chapter 6, are exactly this: causal base rates about groups. That's why they stick when statistics slide off.
The operational rule: expect regression anywhere you selected on an extreme. Not "sometimes," not "in sports" — anywhere the thing that got your attention got it by being an outlier on a noisy measurement. A running list from the places you actually work:
And the single portable habit: before crediting any cause for a change back toward normal — a fix, a pep talk, a reorg, a new vitamin — ask one question first. "How noisy is this measurement?" If the answer is "noisy" and the starting point was extreme, regression has already explained most of the change, and the burden of proof is on the story.
Your best engineer had a stellar quarter; this one is merely fine. Something changed. They've lost motivation since the promotion — or they're interviewing, or the new manager is smothering them. The story-writer needs a villain, and it will have one by end of stand-up. Find the cause, fix the person.
A stellar quarter is skill + a good luck draw — the right tickets, a cooperative dependency, a bug that happened to be shallow. The skill renews next quarter; the luck draw doesn't. A merely-fine follow-up isn't a decline from baseline; it usually is the baseline, seen after one lucky sample.
The transferable habit: before hunting causes for a dip from a peak, ask what the measurement's noise band looks like. Most "declines from a peak" are the peak's fault, not the person's.