Game theory
Why cooperation survives when cheating pays
In any single exchange, taking advantage of someone who trusts you is the better move. Cooperation is everywhere anyway. The reason is repetition, and you can run the model yourself.
Two people can each either cooperate or take advantage of the other. If both cooperate, both do well. If both cheat, both do badly. If one cheats while the other cooperates, the cheat does best of all and the cooperator does worst. Work through it from one person's side and the answer is uncomfortable: whatever the other person does, you are better off cheating.
That is the prisoner's dilemma, and the reason it has held attention for seventy years is not that it is clever. It is that the logic is airtight and the conclusion is obviously wrong. Cooperation is not rare. It is most of what happens. Suppliers deliver before they are paid, colleagues cover for each other, strangers give directions. Something is missing from the model.
What the payoffs actually say
The usual numbers: three points each for mutual cooperation, one point each for mutual defection, and where one defects against a cooperator, five to the defector and nothing to the cooperator. Two conditions make it a dilemma rather than an easy choice. Defecting beats cooperating against either kind of opponent — five beats three, one beats nothing. And mutual cooperation beats taking turns exploiting each other, since three each beats an average of two and a half.
So the individually rational move is defection, and a population of individually rational players scores one point each where they could have had three. The trap is not that anyone is being stupid. Each step is correct. Following all of them lands everyone somewhere worse.
What changes when you play again tomorrow
Play once and there is nothing to protect. Play repeatedly, without a known last round, and today's move sets the terms of tomorrow's. The defection still earns five, and it also buys an opponent who has learned something about you.
In 1980 Robert Axelrod ran this as an open tournament. Game theorists, economists, psychologists and computer scientists submitted strategies, and every strategy played every other over hundreds of rounds.[2] The entries included some genuinely elaborate machinery. The winner was four lines long: cooperate on the first move, then do whatever the opponent did last time. Anatol Rapoport had submitted Tit for Tat. Axelrod published the results, invited a second round with everyone now knowing what had won, and Tit for Tat won again.[1]
The interesting part is how. Tit for Tat never beats anyone. Against any single opponent it either draws or loses — it cannot do better, because it never defects first and never defects more than the other side did. It wins the tournament by never doing badly, while the aggressive strategies win their individual matches and wreck each other in the process.
That is a claim worth testing rather than taking. Below is the tournament, running in this page. Set the rounds, pick who plays, and see the table for yourself.
Every strategy plays every other, including a copy of itself. Change the number of rounds, add noise so moves are occasionally misread, or switch to evolution and let the population reproduce in proportion to score.
Requires KIT Apps · Runs locally · 25 KBKIT Apps is free
The first thing you will notice is that Tit for Tat does not win. Grudger does — the strategy that cooperates until it is crossed once and then defects forever. That is not a bug in the model, and it is not a contradiction of Axelrod either. Against every nice strategy Grudger behaves exactly like Tit for Tat, because it never gets a reason to hold its grudge. Against Random it does better, because it stops being exploited after the first betrayal instead of trading punishments forever. In a world where nobody ever defects by accident, never forgiving costs nothing.
Hold on to that, because it is the setup for the part that matters. Axelrod's tournament was a particular set of entries; the general claim that survived was about the four properties below, and about what happens when the world stops being clean.
Drop the rounds to two or three and Always Defect climbs the table, because there is no future left to protect and the game has collapsed back into the one-shot version. Push the rounds up and it sinks. Nothing about the strategies changed. The length of the relationship changed.
What wins, and what that tells you
Axelrod's reading of the results was that the strategies which did well shared four properties, and the tool makes each of them checkable:
- Nice. Never the first to defect. Every high scorer in the original tournament was nice, and every low scorer was not.
- Retaliatory. Responds to defection. Always Cooperate is nice and finishes badly, because it funds whoever exploits it.
- Forgiving. Goes back to cooperating once the other side does. Compare Tit for Tat with Grudger, which never forgives: one recovers, the other spends the rest of the game in mutual punishment.
- Clear. Easy to work out. An opponent who cannot predict you cannot learn to cooperate with you, which is why Random scores poorly and drags others down with it.
Now make the world realistic
Real exchanges are noisy. A message is missed, an invoice goes to the wrong address, a deadline slips for reasons nobody controlled. The other side sees a defection you did not make.
Turn the noise slider up in the model above and watch what breaks. Two copies of Tit for Tat, both trying to cooperate, hit a single misread move and lock into alternating retaliation — each punishing the other for the punishment before. Neither did anything wrong. Under noise the strategies that hold up are the ones with slack in them: Tit for Two Tats, which waits for a second defection before reacting, and Generous Tit for Tat, which forgives about one defection in ten.[4] Pavlov — repeat your last move if it worked, switch if it did not — also does well here, and under some conditions outperforms Tit for Tat outright.[3]
The practical version of that finding: in any system where signals are imperfect, a policy of immediate proportional retaliation will manufacture conflicts out of accidents. Some deliberate slack is not softness. It is what stops noise from compounding.
And watch Grudger while you do it. The strategy that topped the clean table falls down it as soon as mistakes are possible, because a single accident costs it every remaining round of that relationship. Its advantage was never robustness. It was the absence of accidents. Very few real arrangements have that.
Where this actually applies
The model is not a theory of human nature and it does not need to be. It is a claim about structure: the same people behave differently depending on whether the game repeats, and the structure is usually easier to change than the people.
- A supplier who expects one transaction and a supplier on a rolling contract are playing different games, and will behave differently without either of them changing their character.
- Teams that reorganise constantly keep resetting the round counter. Nobody expects to deal with the same people next quarter, so nobody builds the reputation that makes cooperation pay.
- A supplier's last quarter before an acquisition, a manager's last month in post, a contractor's final invoice: a known last round removes the reason to cooperate in it, and then in the round before it.
The design question is rarely how to make people more trustworthy. It is whether the arrangement lets them find out that cooperating is worth it, and whether it survives the first misunderstanding. Those are both things the model will tell you — set it up the way your situation is set up, and run it.
The short version
Sources
- 1.Axelrod, R. and Hamilton, W. D. (1981). The Evolution of Cooperation. Science, 211(4489), 1390–1396.
- 2.Axelrod, R. (1980). Effective Choice in the Prisoner's Dilemma. Journal of Conflict Resolution, 24(1), 3–25.
- 3.Nowak, M. and Sigmund, K. (1993). A strategy of win-stay, lose-shift that outperforms tit-for-tat in the Prisoner's Dilemma game. Nature, 364, 56–58.
- 4.Nowak, M. and Sigmund, K. (1992). Tit for tat in heterogeneous populations. Nature, 355, 250–253.
- 5.Stanford Encyclopedia of Philosophy: The Prisoner's Dilemma.