Algorithms to Live By · ch.11 · game theory
♟ Chapter 11 · The minds of others

Game theory:
everyone else is thinking too

Every problem so far had one solver — you. Add a second person who's also optimising, and the ground moves: the best move now depends on their best move, which depends on yours.

Ten chapters, one solver. You against a fixed world: unknown restaurants, unsorted shelves, a cache that doesn't care what you want. The world was hard, but it was never trying.

Now put a second optimiser in the room. Everything you've learned still applies — and none of it is enough, because the thing you're reasoning about is reasoning back.

1The recursion that eats itself

You're at a poker table. He pushes his stack in and you think: he's bluffing. That's level one.

Then it lands: he knows I think he's bluffing — so maybe the shove is real, priced to look fake. Level two. And then: he knows I know he knows, so maybe it's a bluff after all. Level three. There is no level at which this stops; each floor is built out of the floor above it.

Any programmer feels the shape of this before they can name it. It's a recursive call with no base case — not a hard problem, a non-terminating one. Your model of him contains a model of you, which contains a model of him, and the stack grows until something gives. The something is usually you.

Game theory's founding move is to refuse the recursion entirely. Stop asking "what will they do?" — that question is the loop. Ask instead: where does this settle? Don't trace the calls. Find the fixed point.

🌀
The switch is the same one you make when you stop unrolling a recursive definition by hand and start looking for the value that satisfies it. You never simulate f(f(f(...))). You solve for x = f(x) and move on with your life.

2Equilibrium: the place nobody wants to leave

Here's the base case. Look for a set of choices — one per player — where nobody can do better by changing their mind on their own. If you switch, you lose. If they switch, they lose. So nobody switches. The recursion terminates not because someone out-thought someone else, but because there's nowhere left to go.

Rock-paper-scissors is the clean demo. Play each option exactly a third of the time, at random, and there is literally nothing your opponent can do about it. Rock every time? You win a third, lose a third, tie a third. A brilliant read of your tells? Same. Your strategy has made theirs irrelevant — which means it's made all that reading worthless. You can play rock-paper-scissors optimally against a mind reader.

That resting point is a Nash equilibrium, and in 1951 John Nash proved something that sounds modest and isn't: every finite game has at least one. Any game, any number of players, any payoffs — allow mixed (randomised) strategies and a settling point is always there.

That proof is what makes the infinite regress stop. You don't have to out-think anyone to the eleventh level, because the eleventh level converges to the same place the third one did. There is always a floor.

Notice what an equilibrium is not. It isn't a prediction that people are clever, it isn't a plan, and — hold this thought — it is emphatically not a promise that the outcome is any good.

3Finding it is another matter

Nash tells you the equilibrium is in the building. He does not tell you which room.

His is an existence proof — the mathematical equivalent of "there's definitely a solution, good luck" — and for decades everyone assumed finding it was a mopping-up exercise. It wasn't. In 2006 Daskalakis, Goldberg and Papadimitriou showed that computing a Nash equilibrium is PPAD-complete: intractable, in the same defeated tone of voice as NP-hard.

This has teeth, and they bite in an unexpected direction. If nobody — not a person, not a market, not a data centre — can compute where a game settles, then "where it settles" is a bad prediction of what people will actually do. An equilibrium too expensive to find isn't a description of behaviour; it's a fact about a mathematical object that human beings never visit. As Papadimitriou put it: if your equilibrium concept isn't efficiently computable, it's a poor model of a market.

Which is the Chapter 8 lesson wearing a different hat. When the exact answer is out of reach, the exact answer stops being the right goal. Relax it, approximate it, or ask a different question — and notice that every practical result in the rest of this chapter comes from asking a different question.

4Equilibrium ≠ good

Two people are arrested and put in separate rooms. Each is offered the same deal: rat out the other and walk, while your partner takes the long sentence. If you both rat, you both do medium time. If you both keep quiet, the case is thin — short stretch, then home.

Sit in one of those rooms and do the arithmetic. Suppose they stay quiet: ratting gets you out today instead of next month. Rat. Suppose they rat: ratting gets you medium instead of maximum. Rat. You never needed to know what they'd do. Ratting is better in every single case.

A move that's better no matter what the other person does is a dominant strategy, and finding one feels like winning — no reading, no recursion, no guessing. Except the person in the other room is running the identical arithmetic and reaching the identical conclusion, and you both do medium time in a world where you could both have gone home in a month.

Everyone was rational. Everyone was right. And the equilibrium — the only place neither of you wants to move from — is worse for both of you than a place you could both see, could both name, and could not reach.

That's the chill of it. The prisoner's dilemma isn't a story about bad people, or weak people, or people who didn't think it through. It's a story about people who thought it through perfectly and arrived at the bad place on purpose, one flawless step at a time.

Interactive · the dilemma, and how to break it Edit the payoffs · watch the equilibrium move
Rounds
0
Your score
0
Their score
0
Per round

Play it once and defecting is unanswerable. Play it over and over against the same person and the arithmetic changes, because today's move buys tomorrow's reply. In 1980 Robert Axelrod ran exactly that as a tournament, inviting game theorists to submit programs. The winner was the shortest entry: Tit-for-Tat. Open with cooperation, then copy whatever they did last time. Four lines.

Run the tournament above and watch the thing everyone misses. Tit-for-Tat never beats a single opponent head-to-head — it can't, by construction; it only ever mirrors, so a draw is its ceiling against anyone who isn't flipping coins. It wins the whole tournament while drawing or losing every match in it. Its trick isn't out-scoring you. It's refusing to be exploited while making it cheap for you to be decent, and collecting on that across every game at once.

🤝
Now push the payoff steppers. Drag the temptation below the reward — make ratting pay less than staying quiet — and watch the coral equilibrium marker jump to the cooperate corner and the dilemma simply evaporate. The players didn't change. Nobody had a talk about values. The numbers changed, and the trap wasn't there any more. Hold on to that; section 6 is built on it.

5The price of anarchy

"Everyone acting selfishly ends up worse" is a bar-stool observation. The useful question is how much worse. Take the outcome when everybody optimises for themselves, take the outcome a benevolent dictator would have arranged, and divide. One number. If it's 1.0, selfishness costs nothing at all. If it's 4, the free-for-all is four times as bad as it needed to be.

That ratio is the price of anarchy, and it turns a moral argument into a measurement — which is the point, because now you can be surprised by the answer instead of confirming what you already believed.

It bites hardest in traffic. Every driver picks the route that's fastest for them, given what everyone else is doing. Perfectly sensible, perfectly selfish, and the result is a real equilibrium: nobody can improve by rerouting alone. Then comes the twist that ought to be impossible.

Add a new road. A good one — fast, free, exactly where people said they wanted it. And every commute gets longer. Not from construction, not from induced demand over years. Immediately, at equilibrium, on the arithmetic. This is Braess's paradox, and the machine below runs it live.

Interactive · the road that makes traffic worse Equilibrium solved numerically each frame
4000
Avg commute now
Without the road
Social optimum
Price of anarchy

Watch what happens when you open it. The shortcut is free and instant, so from any driver's seat it is strictly better. Every driver reasons correctly. Every driver takes it. And the routes that used to soak up half the traffic each now sit empty while everyone piles onto the two congestible roads at once. Nobody can fix it alone: refusing the shortcut just makes your commute worse while everyone else keeps theirs. The trap is airtight, and it was built by a road nobody had to use.

The half of the story that gets skipped is the good half. Selfish routing has a price of anarchy of about 4/3 — a hard ceiling, proven by Roughgarden and Tardos. In the worst case ever, over any network, any demand, everyone-for-themselves is 33% worse than a perfectly coordinated dictator. That's the disaster scenario, and it's… fine? Decentralised traffic is nearly optimal. Leave it alone. A control tower routing every car would cost more than the 33% it recovers.

Compare the prisoner's dilemma, where the price of anarchy is unbounded — you can make the gap between the equilibrium and the deal as catastrophic as you like just by picking the payoffs. Same word, "selfishness"; two completely different situations. One deserves a shrug, one deserves an intervention.

📏
Measure the price before you moralise about the players. "People are being selfish" is not a finding — selfishness is a constant. The finding is the ratio. Low price of anarchy? Decentralisation is working; your urge to add process is the actual problem. High? Then no amount of exhortation will help, and you need to change the structure.

6Change the game, not the players

Everyone on the team is answering Slack at eleven at night, everyone is exhausted, and everyone would rather not. So somebody sends a message about work-life balance and boundaries. Nothing changes — and the confusing part is that nothing changes even though everyone agreed.

It failed because it was aimed at the wrong layer. Nobody is answering Slack at eleven because they forgot they shouldn't. They're answering because the colleague who answers looks committed, and the payoffs of looking committed are exactly what you'd expect. They are playing the game correctly. You just don't like the game.

Which makes the entire genre of "be better" advice a category error. You cannot fix a bad equilibrium from inside it — that is what the word equilibrium means. Unilateral improvement is impossible, by definition. Asking people to unilaterally improve is asking them to take the sucker's payoff and call it growth.

  • The commons. Everyone's cows overgraze the shared field until there's no field. Every herder correctly notices that their cow is a rounding error and the grass is going anyway. Nobody is villainous. The field still dies.
  • Unlimited vacation. Remove the cap and people reliably take less time off. A fixed 25 days is a floor you're a fool not to use — it's yours, it expires, take it. "Unlimited" deletes the floor and replaces it with a comparison to your colleagues, and the equilibrium of that comparison is a race to look indispensable. Generosity, arranged so it can't be accepted.
  • The arms race. Anywhere the only reward is being ahead — hours worked, arriving early, escalating armament — the equilibrium is maximum effort and no benefit to anyone, because being ahead is zero-sum and the effort isn't.

Each dissolves the instant somebody changes the numbers instead of the sermon. Lights out at six. A hard 25-day minimum, taken. A treaty, a union contract, a rule everyone hates and everyone quietly relies on. Not because a rule makes people good — because a rule lowers the temptation payoff, and once ratting stops paying, nobody has to be a hero to stop ratting.

A rule beats willpower every time, because willpower is a unilateral deviation and rules move the equilibrium.

This is the inverse problem, and it has a name: mechanism design. Ordinary game theory asks "given these rules, what happens?" Mechanism design runs it backwards — "given the outcome I want, what rules produce it?" — and it's the half of the field you can actually use on a Tuesday. You are rarely the player. You are surprisingly often the person who gets to write the rules.

7The Vickrey auction — honesty as an equilibrium

Now the beautiful one.

Sealed-bid auction, ordinary rules: everyone writes a number, highest bid wins and pays what they wrote. Think about how you'd play. The thing is worth 100 to you, so bidding 100 is pointless — you'd win and profit nothing. Bid 70? Only if you think the others are under 70. You're back at the poker table, modelling their model of you, and every bid is regrettable in two directions at once: too high and you overpaid, too low and you lost something you wanted. You'll never learn which mistake you made, because you only ever observe the outcome you caused.

Now change one word. Highest bid still wins — but the winner pays the second-highest bid.

Play it out. Your bid no longer sets your price; it only sets whether you win. So bid what the thing is honestly worth to you. Bid lower and you sometimes lose auctions you'd have profited from — the price was below your value anyway, so that was free money you passed up. Bid higher and you sometimes win at a price above your value, which is a loss you signed up for. Both directions are strictly worse. Bidding your true value isn't merely allowed, it's dominant: it beats every alternative no matter what anyone else does.

Sit with what that costs. No modelling the room. No recursion. No shading, no bluffing, no regret. The mechanism does the thinking, and what it hands back is the instruction "just say what it's worth."

Interactive · expected profit vs. your bid Rivals' values ~ Uniform(0,100) · exact expectations
80
80
3
First-price sealed bid Second-price (Vickrey) English / open outcry ★ = best bid for that format
First-price profit
Vickrey profit
Win probability
Vickrey's best bid

Drag your bid across the green curve and try to find a way to profit by lying. There isn't one — the peak sits at your true value exactly, for every value and every rival count you can dial in. Now look at the blue first-price curve, whose peak has wandered off below your value to a spot you could only find if you already knew how many rivals you had. That gap between the peak and the truth is the strategising, drawn to scale.

The coral dashes are the English auction — the open, shouting, paddle-raising one — landing precisely on top of the Vickrey curve. Not a coincidence, and the whole reason the format survives: raising your hand until the price passes what it's worth to you and then stopping is bidding your true value, and the winner pays roughly the runner-up's walk-away point. The theatre is a slow reveal of a sealed second-price auction.

A mechanism where honesty is your best move no matter what anyone else does is strategy-proof — or, in the field's own phrasing, incentive-compatible. Vickrey won a Nobel for it. It runs ad auctions, spectrum sales, and the kidney exchanges that match donors to strangers, where "just tell us the truth" isn't a nicety but the only way the thing works at all.

💎
This is the chapter's one piece of good news, and it's worth being precise about why. Vickrey didn't make bidders honest. He made honesty selfish — the move you'd choose anyway, out of naked self-interest, having considered the alternatives and found them worse. That's not a moral achievement. It's an engineering one, and unlike a moral achievement you can ship it and it stays shipped.

8Information cascades and the madness of sensible people

Two restaurants, side by side, both unknown to you. One has a queue. You join it — obviously. Those people know something. So does the next person, who sees a slightly longer queue and joins for the same excellent reason. By the fiftieth person the queue is overwhelming evidence, and here is what it is evidence of: two people, at the start, who had a hunch.

Nobody was foolish. Every person in that line performed a correct inference from the information available. The failure is structural: after the first couple of guesses, everyone is reading the guesses and not the evidence, and the guesses are all downstream of the same two hunches. The chain is long and anchored to nothing. That's an information cascade, and once you know the shape you find it under every bubble — tulips, dot-coms, houses — because a price is a queue that reports itself as a fact.

Keynes found it in the markets and called it the newspaper beauty contest: readers pick the prettiest face from a set of photos, and the prize goes to whoever picks the most popular face. So you don't pick who you find beautiful. You pick who you think others will pick — except they're doing the same, so you pick who you think others think others will pick. There's the poker table again, the same recursion all the way down, the actual faces long since irrelevant. Everyone trading on a consensus that nobody holds and nobody has checked.

The practical advice is embarrassingly small: say what you actually think, not what you think the room thinks. When you catch yourself quietly revising toward the mood of a meeting, notice that your private information is the only thing you brought that isn't already in the room. Launder it into an echo and the group gets nothing while the cascade gets one more link. Speak first if you can; if you can't, speak anyway.

📣
The inverse, if you're running the meeting: you are the mechanism. Ask for opinions in writing, at the same time, before anyone speaks. You have just deleted the cascade — not by asking people to be braver, but by removing the sequence that made deference rational. Change the game, not the players.

9The tournament is a fold

The iterated prisoner's dilemma is one of those problems whose type signature does most of the explaining. A Strategy is a function from history to a move — that's the whole interface, and every player above fits it. A match is a fold that carries the history and two scores; a tournament is a fold over matches. Once Strategy is a plain function, Tit-for-Tat stops being a theory and becomes a line of code you could have written by accident.

💡
Read titForTat again — maybe C snd (listToMaybe history) — and notice it's three ideas, not three lines. Be nice: with no history, cooperate. Retaliate: otherwise, do what they just did, so defection costs them immediately. Forgive: look only at the last round, so one cooperation from them and it's over, no grudge. Delete the forgiveness and you have grim, which never rebuilds anything it loses. The difference between the winning strategy and the bitter one is the length of the history it consults.

10Check yourself

3 questions · instant feedback 0 / 3

11What's coming

This chapter kept arriving at the same place from different directions: when the outcome is bad, the lever is the game, not the goodwill. One chapter left, and it takes that lever and points it somewhere unexpected — not at markets or traffic or auctions, but at the ordinary business of dealing with each other. "Wherever you like, I don't mind" sounds generous and is a computational assault. The finale is about the design problem hidden in every interaction: making other people's problems easier.

💝
Up next · Chapter 12
Computational Kindness
The finale. Why "whatever you want" is cruelty in disguise, why constraints are a gift, and how to hand people problems they can actually solve — mechanism design pointed at the person across the table.