Solutions to the example questions on the Nash Equilibrium page.
(a) Symmetric zero-sum game. [3]
The column player's matrix is \(M_c = -M = \begin{pmatrix} 0 & 1 & -1 \\ -1 & 0 & 1 \\ 1 & -1 & 0 \end{pmatrix}\). It is zero sum because \(M + M_c = 0\): one player's gain is exactly the other's loss. It is symmetric because \(M = -M^{\top}\), so interchanging the two players maps the game to itself while negating the payoffs.
(b) Utilities for the three strategy pairs. [3]
The row player's utility is \(u_r = \sigma_r M \sigma_c^{\top}\) and, as the game is zero sum, the column player's is \(u_c = -u_r\). For \(\sigma_c = (\tfrac{1}{2}, \tfrac{1}{2}, 0)\) we have \(M \sigma_c^{\top} = (-\tfrac{1}{2}, \tfrac{1}{2}, 0)\).
(c) The uniform strategy is a Nash equilibrium. [5]
The best response condition states that \(\sigma_r\) is a best response to \(\sigma_c\) if and only if every action in its support yields the maximal expected payoff, that is \(\sigma_r(i) > 0 \Rightarrow (M\sigma_c^{\top})_i = \max_k (M\sigma_c^{\top})_k\). Take both players uniform, \(\sigma = \left(\tfrac{1}{3}, \tfrac{1}{3}, \tfrac{1}{3}\right)\). Against \(\sigma_c = \sigma\), the row player's expected payoffs are
All three actions give the maximal payoff \(0\), so every action in the support of \(\sigma_r\) is a best response: by the best response condition \(\sigma_r\) is a best response. By the symmetry of the game the same holds for the column player, so \((\sigma, \sigma)\) is a Nash equilibrium, with value \(0\).
(d) Uniqueness. [5]
In Rock-Paper-Scissors every action is beaten by another, so no Nash equilibrium can place all weight on a proper subset of the actions: whatever the support of the opponent's strategy, a player can deviate to the action that beats the opponent's most likely choice. Hence in any equilibrium each player mixes over all three actions, and by the best response condition is indifferent between them. Writing \(\sigma_c = (a, b, c)\) with \(a + b + c = 1\), the row player's payoffs are
Setting them equal, \(c - b = a - c\) gives \(2c = a + b\), and \(a - c = b - a\) gives \(2a = b + c\); with \(a + b + c = 1\) these force \(a = b = c = \tfrac{1}{3}\). By symmetry \(\sigma_r\) is also uniform, so the equilibrium is unique.
(e) Rock-Paper-Scissors-Lizard-Spock. [3]
By the same symmetry each action beats two others and loses to two, so against the uniform strategy each action earns \(\tfrac{2(+1) + 2(-1)}{5} = 0\): all five are best responses. As in part (d), no equilibrium can leave any action unused, so the unique Nash equilibrium is each player playing each of the five actions with probability \(\tfrac{1}{5}\), with value \(0\).
(a) Definitions. (Bookwork.) [5]
(b)(i) Pure Nash equilibria. [4]
Underlining best responses (each column for the row player in \(M_r\), each row for the column player in \(M_c\)):
Both payoffs are underlined in the top-left and bottom-right cells, giving the pure Nash equilibria \(\{((1, 0), (1, 0)), ((0, 1), (0, 1))\}\).
(b)(ii) Sketches and indifference. [6]
Against \(\sigma_2 = (y, 1 - y)\): \(u_1(r_1) = 4y + (1 - y) = 1 + 3y\) (from \((0,1)\) to \((1,4)\)) and \(u_1(r_2) = 3(1 - y) = 3 - 3y\) (from \((0,3)\) to \((1,0)\)); they cross where \(1 + 3y = 3 - 3y\), i.e. \(y = \tfrac{1}{3}\). Against \(\sigma_1 = (x, 1 - x)\): \(u_2(c_1) = 3x + (1 - x) = 1 + 2x\) and \(u_2(c_2) = 4(1 - x) = 4 - 4x\); they cross where \(1 + 2x = 4 - 4x\), i.e. \(x = \tfrac{1}{2}\).
(b)(iii) Mixed Nash equilibrium. [5]
By the best response condition, writing the row player's matrix as \(A = M_r\), a strategy \(\sigma_r^{*}\) is a best response to \(\sigma_c\) if and only if
that is, every action played with positive probability must earn the maximal expected payoff. For an equilibrium in which both supports have size two, every action lies in the support, so each player must be indifferent between their two actions and hence make the other indifferent. From part (ii) the column player is indifferent at \(x = \tfrac{1}{2}\) and the row player at \(y = \tfrac{1}{3}\), so the mixed Nash equilibrium is \(\left(\left(\tfrac{1}{2}, \tfrac{1}{2}\right), \left(\tfrac{1}{3}, \tfrac{2}{3}\right)\right)\).
(b)(iv) Payoffs and comparison. [4]
At the mixed equilibrium each player is indifferent, so the row player earns \(u_1(r_1) = 1 + 3 \cdot \tfrac{1}{3} = 2\) and the column player earns \(u_2(c_1) = 1 + 2 \cdot \tfrac{1}{2} = 2\); the payoff profile is \((2, 2)\). The pure equilibria give \((4, 3)\) at \(((1,0),(1,0))\) and \((3, 4)\) at \(((0,1),(0,1))\). Each player earns at least \(3\) in either pure equilibrium, which exceeds the mixed payoff of \(2\), so both players would prefer either pure equilibrium to the mixed one: the mixed equilibrium is the worst of the three for both players.
(a) Definitions. (Bookwork.) [5]
(b)(i) Iterated elimination. [4]
For the row player, row 2 strictly dominates row 1 (\(5 > 3\) and \(1 > 0\)), so row 1 is eliminated. With only row 2 remaining, the column player's payoffs are \(0\) for column 1 and \(1\) for column 2, so column 2 strictly dominates column 1 and column 1 is eliminated. The single surviving profile is \(((0, 1), (0, 1))\), the Nash equilibrium, with payoffs \((1, 1)\).
(b)(ii) Interpretation. [3]
The outcome in which each plays their first action gives \((3, 3)\), which both players prefer to the equilibrium \((1, 1)\) since \(3 > 1\). It cannot be sustained in a one-shot game: from \((3, 3)\) either player can deviate to their second action and earn \(5 > 3\), so it is not a Nash equilibrium. Strict dominance drives both players to their second action and hence to the inferior outcome \((1, 1)\).
(c)(i) Pure Nash equilibria. [3]
The pure Nash equilibria are \(\{((1, 0), (1, 0)), ((0, 1), (0, 1))\}\), with payoffs \((2, 1)\) and \((1, 2)\).
(c)(ii) Mixed Nash equilibrium. [5]
By the best response condition each player makes the other indifferent. The row player makes the column player indifferent: \(u_2(c_1) = x\) and \(u_2(c_2) = 2(1 - x)\), equal at \(x = 2 - 2x\), so \(x = \tfrac{2}{3}\). The column player makes the row player indifferent: \(u_1(r_1) = 2y\) and \(u_1(r_2) = 1 - y\), equal at \(2y = 1 - y\), so \(y = \tfrac{1}{3}\). The mixed Nash equilibrium is \(\left(\left(\tfrac{2}{3}, \tfrac{1}{3}\right), \left(\tfrac{1}{3}, \tfrac{2}{3}\right)\right)\).
(c)(iii) Payoffs and inefficiency. [3]
At the mixed equilibrium the row player earns \(u_1(r_1) = 2y = \tfrac{2}{3}\) and the column player earns \(u_2(c_1) = x = \tfrac{2}{3}\). Each pure equilibrium gives the players \(2\) and \(1\) (in some order), and even the smaller of these, \(1\), exceeds \(\tfrac{2}{3}\). So the mixed equilibrium is worse for both players than either pure equilibrium: a failure to coordinate.
(a) Antisymmetry forces a zero diagonal value. [5]
The quantity \(s = \sigma A \sigma^T\) is a scalar, so it equals its own transpose. Using \(A^T = -A\),
Hence \(s = -s\), so \(s = 0\). A player who uses any strategy \(\sigma\) against an opponent using the same \(\sigma\) therefore earns an expected payoff of zero.
(b) Equilibrium condition and value. [6]
By the best response condition, \(\sigma^*\) is a best response to \(\sigma^*\) if and only if every action in the support of \(\sigma^*\) is itself a best response, and no action does strictly better. The payoff to playing pure action \(i\) against \(\sigma^*\) is \((A \sigma^{*T})_i\), while \(\sigma^*\) itself earns \(\sigma^* A \sigma^{*T} = 0\) by part (a). So \((\sigma^*, \sigma^*)\) is a Nash equilibrium if and only if
Playing \(\sigma^*\) guarantees the row player an expected payoff of \(0\) against any opponent strategy, and by symmetry the column player can likewise guarantee \(0\); since the game is zero-sum the two guarantees are consistent only if the value is exactly \(0\).
(c) The weighted equilibrium. [8]
A full-support equilibrium has \(A \sigma^{*T} = 0\). Writing \(\sigma^* = (x, y, z)\),
which are consistent with \(x = y = \beta z\). The normalisation \(x + y + z = 1\) gives \(z(2\beta + 1) = 1\), so
When \(\beta = 1\) this is \(\bigl(\tfrac{1}{3}, \tfrac{1}{3}, \tfrac{1}{3}\bigr)\), recovering standard Rock-Paper-Scissors.
(d) Interpretation. [6]
As \(\beta\) increases the equilibrium weight on Rock and Paper rises towards \(\tfrac{1}{2}\) each while Scissors falls towards \(0\); as \(\beta \to 0\) almost all the weight goes onto Scissors. The parameter \(\beta\) scales the stakes of the contests that Scissors is involved in: a larger \(\beta\) makes the swings around Scissors more severe, and in equilibrium the players protect themselves by playing Scissors less often and the other two actions more. The equilibrium always keeps every action a best response, so no action is ever abandoned for any finite \(\beta > 0\), but the mixing tilts smoothly with the weight.
Marking exercise 1 (Question 2(b)(iii)).
The transcript contains three errors.
Step 1 is also looser than the best response condition: indifference is required across the actions in the support of the strategy. It does no harm here because the equilibrium has full support, but an examiner expects the condition stated precisely. A fair mark is [2] of [5]: the structure and the arithmetic earn credit, but the strategy pair reported is not an equilibrium of the game and the uniqueness claim is false.
Marking exercise 2 (Question 3(b)(i) and (ii)).
For part (i), the eliminations and the surviving cell are correct, but the domination is strict, not weak: \(5 > 3\) and \(1 > 0\) are strict inequalities in both comparisons. The distinction carries weight, since iterated elimination of strictly dominated strategies can never remove a Nash equilibrium, which is exactly why the argument identifies the equilibrium; a weakly dominated action can feature in an equilibrium. A fair mark is [3] of [4].
For part (ii), the claim that \((1, 1)\) is Pareto efficient is false: \((3, 3)\) makes both players strictly better off, so \((1, 1)\) is Pareto dominated, and this inefficiency is the entire point of the Prisoner's Dilemma. The property the transcript describes, that no player gains by changing their own strategy alone, is the definition of a Nash equilibrium, not of Pareto efficiency: the two concepts have been swapped. The final sentence on the incentive to deviate is correct. A fair mark is [1] of [3].