Skip to the main content

Experiment 01 Repeated Games

The Trust Machine

Will you cooperate when betrayal pays more today? Play repeated rounds and find out what tomorrow changes.

5–10 minutes to explore Prototype Updated

Blue and orange playing pieces stand beside two matching halves of a shared bridge.
About these models Math step by step For experts Math symbol guide

Loading the interactive experiment…

Start with “Work It Out, Step by Step” below. The optional expert section explains its symbols as you go. For more examples, use the plain-language math guide.

1. Make a Prediction

Can Trust Survive a Mistake?

You and another player choose at the same time. Cooperate means help make a good result for both of you. Defect means take the other choice, which pays you more in this round if the other player’s choice stays the same. A strategy is a rule for deciding what to do.

If both cooperate, each gets 3 points. If only one defects, that player gets 5 and the player who cooperated gets 0. If both defect, each gets 1.

2. Make a Consequential Choice

Meet Your Opponent

Choose an action to play your first round.

Your encounter: actual actions and points in each round
RoundYouOpponentYour pointsTheir pointsCombined

3. Change One Assumption

Set the Conditions

Changing a setting starts a fresh encounter. The seed replays the same random sequence when you make the same choices. Shared links preserve these settings, not your play history.

With a known final round, the continuation percentage is unused. With a random ending, the known round count is unused; an encounter is capped at 200 rounds for responsiveness.

Compare Strategies Across Many Encounters

Each of the six strategies plays every strategy, including another copy of itself, the same number of times. The ranking uses average points per round: total points divided by rounds played. When a strategy plays its own copy, we average the two copies so that pairing does not count twice. “Cooperation rate” means the percentage of actions that were cooperate. Run again with a different seed to see whether the result changes.

Simulated tournament results
StrategyPoints per roundCooperation ratePlayer-rounds

Reveal the Incentives

Work It Out, Step by Step

Worked example using the default settings: 10 rounds, no mistakes, and a tit-for-tat opponent. These numbers explain the starting game; they do not update when you change the controls.

Tit for tat means “cooperate first, then copy the other player’s last action.” You only need addition, subtraction, multiplication, and division to compare these choices.

  1. Both cooperate in round 1. You get 3 points and your opponent gets 3. Together, you earn \(3 + 3 = 6\) points.
  2. Keep cooperating for all 10 rounds. Your opponent keeps copying cooperation. Each player earns \(10 \times 3 = 30\) points. Your average is \(30 \div 10 = 3\) points per round.
  3. Now try a different choice in round 1. If you defect while your opponent cooperates, you get 5 points instead of 3. That is \(5 - 3 = 2\) extra points for you. Your opponent gets 0.
  4. Check what happens next. In round 2, tit for tat copies your defection. If you now cooperate, you get 0 points. Your total for these two rounds is \(5 + 0 = 5\). Cooperating in both rounds would have earned you \(3 + 3 = 6\).

A choice that wins more points now can change what happens later. This two-round comparison does not prove that one strategy always wins. Try a known final round, mistakes, or a different opponent to see why the conditions matter.

If you turn on mistakes, a 10% chance means about 10 flips per 100 actions over many trials, not exactly 10 every time. An accidental defection can start a chain of responses even when neither player meant to break cooperation.

For experts: formal model and assumptions

How to Read the Symbols

The symbols below are short names for points, rounds, and chances. You can use the game without memorizing them. The math reading guide has more examples.

Letters and small labels
A letter stands for a value. A small letter below it is a label: \(u_i\), read “u for player i,” means that player’s points. The letter \(i\) identifies a player; it is not a number to multiply by. Lowercase \(u\) means points in one round, uppercase \(U\) means a total, and a bar above \(u\) means an average. Read about letters and labels.
Inputs in parentheses
\(u_i(t)\), read “u for player i at round t,” asks for that player’s points in round \(t\). Here the parentheses identify the round; they do not mean multiplication. Read about functions.
Comparison and multiplication
\(>\) means “greater than,” and \(\le\) means “less than or equal to.” Letters or numbers written next to each other usually mean multiply: \(2R\) means 2 times \(R\), and \(w_kU_k\) means weight times total points. Read about comparisons.
Addition, division, and powers
\(\sum\), the capital Greek letter sigma, means “add these terms.” The labels below and above it say where to start and stop. A fraction bar means divide. A raised number is a power: \(p^2\) means \(p\) times \(p\). The first term in the round-count sum below is 1 because round 1 always happens, even if the chance of continuing is zero. Read about sums and powers.
Chance and expected value
A probability is a chance written from 0 to 1: 0.9 means 90%. \(\mathbb{E}[L]\), read “expected L,” is the average number of rounds the chances predict over many encounters. The square brackets identify what we are averaging; they do not mean multiply. \(\varepsilon\), read “epsilon,” is the Greek letter used here for the chance of an accidental action flip. Read about probability and averages.

Repeated Prisoner’s Dilemma

Let \(C\) denote cooperation and \(D\) defection. A payoff is the points awarded to a player. The stage-game payoffs are \(R=3\) for mutual cooperation, \(S=0\) for cooperating against defection, \(T=5\) for defecting against cooperation, and \(P=1\) for mutual defection. Each cell below lists (row payoff, column payoff).

\[ \begin{array}{c|cc} & C & D \\ \hline C & (R,R)=(3,3) & (S,T)=(0,5) \\ D & (T,S)=(5,0) & (P,P)=(1,1) \end{array} \]

Read the table by finding your row and the other player’s column. In \((3,3)\), the comma separates the row player’s points from the column player’s points; it does not ask you to do arithmetic.

Thus \(T > R > P > S\) reads “5 is greater than 3, which is greater than 1, which is greater than 0.” So \(D\) strictly dominates \(C\) in the stage game: defecting pays more whichever choice the other player makes. Also \(2R > T+S\) says “two times 3 is greater than 5 plus 0.” An encounter adds the points actually earned, with a point in a later round counting just as much as one now. For player \(i\), let \(u_i(t)\) be the points in round \(t\), and let \(L\) be the number of rounds actually played:

\[ U_i = \sum_{t=1}^{L} u_i(t), \qquad \bar u_i = \frac{U_i}{L}. \]

Read it aloud: “Total points for player i equal the sum of that player’s points from round 1 through round L. Average points equal total points divided by rounds.” For 10 rounds earning 3 each, this means add ten 3s to get 30, then divide by 10 to get 3.

A “horizon” is how long an encounter lasts. For a known horizon, \(L=H\), where \(H\) is the selected round count. With continuation probability \(p\), the first round always occurs and each later round requires a fresh draw to continue. For example, \(p=0.9\) means a 90% chance of another round after each round. The expected round count is \(1/(1-p)\), meaning “subtract p from 1, then divide 1 by that result,” without a cap; the implemented cap gives:

\[ \mathbb{E}[L]=\sum_{t=1}^{200}p^{t-1} =\frac{1-p^{200}}{1-p}, \qquad 0\le p\le 0.95. \]

Read it aloud: “The expected number of rounds is the sum of the chances of reaching rounds 1 through 200.” In the sum, \(t\) counts the round, so the first three terms are \(p^0\), \(p^1\), and \(p^2\). At a 90% continuation chance, those are 1, 0.9, and 0.81. The fraction is a shortcut for adding all 200 terms; \(p^{200}\) means multiply 200 copies of \(p\). The comparison at the end says that \(p\) can be anywhere from 0 through 0.95, including both ends.

Each intended action flips independently with mistake probability \(\varepsilon\). Both policies act from completed rounds, then observe actual actions without learning whether a flip occurred. Tit for tat begins with \(C\) and copies the opponent’s last actual action; its forgiving variant chooses \(C\) after \(D\) with the selected probability \(f\). The coin policy independently chooses \(C\) with probability \(1/2\). The final-round variant chooses \(D\) on round \(H\) only with a known horizon.

Tournament ranking pools points and rounds rather than averaging encounter averages. For each strategy, each player appearance \(k\) contributes points \(U_k\) and rounds \(L_k\), weighted by \(w_k=1\) in cross-strategy matches and \(w_k=1/2\) for each copy in a self-match:

\[ \text{tournament score}=\frac{\sum_k w_k U_k}{\sum_k w_k L_k}. \]

Read it aloud: “Add the weighted point totals, add the weighted round counts, then divide the first sum by the second.” Here \(k\) labels each player appearance in the tournament, and \(\sum_k\) means add across all those appearances. \(U_k\) is its point total, \(L_k\) is its round count, and \(w_k\) is its weight. A weight of \(1/2\) counts half of that copy’s points and rounds, so two copies in a self-match together count as one pairing.

Every unordered pairing runs the selected number of trials. Longer encounters contribute more rounds to the pooled score. These fixed-policy simulations neither solve for equilibrium nor predict human behavior. In the separate theoretical game with no mistakes, a commonly known finite horizon, payoff maximization, and common knowledge of rationality, backward induction yields defection every round. Playing against a fixed cooperative policy has different assumptions. Rankings depend on the policy population, noise, horizon, scoring, and sampled random sequence.

Primary reading: Robert Axelrod and William Hamilton, The Evolution of Cooperation (1981). This six-strategy tournament illustrates repeated interaction; it does not reproduce their tournament.

4. Transfer the Lesson

Two teams share an unreliable service. After one missed handoff, one team stops helping. Would punishing every failure forever necessarily improve future cooperation?

Reveal a possible answer

No. If missed handoffs can be accidents, permanent retaliation can destroy useful cooperation. A forgiving response may help when teams expect to work together again, but unconditional forgiveness can also invite exploitation. Investigate the error rate, future relationship, and incentives before choosing a rule.

A Model Is a Place to Start

These small models make the incentives visible. Their results follow from their stated rules; they are not forecasts of how every person or organization behaves. A simulated strategy is a rule, not a personality.

Scenario links save the controls and random seed. To reproduce an interactive run, make the same choices in the same order. Changing a setting restarts the experiment.