Showing posts with label Sequential games. Show all posts
Showing posts with label Sequential games. Show all posts

Thursday, 13 March 2025

Hawks, doves, Israel and Iran

In The Conversation last October, Andrew Thomas (Deakin University) discussed the recent (at that time) military flare-up between Iran and Israel, likening it to a 'game of chicken':

Israel’s strike on military targets in Iran over the weekend is becoming a more routine occurrence in the decades-long rivalry between the two states...

There is a reason why direct military strikes between nations are rare, even between sworn enemies. When attacking another state, it is difficult to know exactly how they will respond, though a retaliatory strike is almost often expected.

This is because defence forces are not just used for fighting and winning wars – they are also vital to deterring them. When a fighting force is attacked, it’s important for it to strike back to maintain the perception it can deter future attacks and make a display of its capabilities. This is what is happening right now between Israel and Iran – neither side wants to appear weak.

If this is the case, where does the escalation end? De-escalation is essentially a game of chicken – one side has to be content with not responding to an attack to take the temperature down.

My ECONS101 class has been covering game theory this week, including the chicken game. In the traditional game of chicken there are two rivals in cars, one at each end of the same street. They drive towards each other at top speed, and whichever rival swerves away first loses the game. So, each rival can choose to speed ahead or swerve away, and each would prefer to speed ahead and win the game. However, the problem is that if both simply keep speeding ahead, it will end in a disastrous crash.

I also recently read the book Hidden Games, by Moshe Hoffman and Erez Yoeli (which I reviewed here). Hoffman and Yoeli have an interesting section in the book on the hawk-dove game, which essentially the chicken game but with a slightly different motivating context. In the hawk-dove game, two rivals are competing over some resource. Each rival can choose to be aggressive or submissive, and whichever rival is more aggressive will win the resource. Each rival would prefer to be aggressive and win the resource. However, if both are aggressive, it ends in a massively disastrous battle.

Coming back to the case of Iran and Israel, this is clearly an example of the hawk-dove game (or the chicken game, if you prefer). This game is laid out in the payoff table below, where the strategies for Israel and Iran are to be aggressive, or submissive. The payoffs are expressed as "+" for good outcomes, and "-" for bad outcomes (and "--" is particularly bad), while zero is a neutral payoff.

To find the Nash equilibrium in this game, we use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium). In this game, the best responses are:

  1. If Iran chooses to be aggressive, Israel's best response is to be submissive (since "-" is better than "--" as a payoff - in other words, taking a bit of punishment is better than a massively disastrous war) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If Iran chooses to be submissive, Israel's best response is to be aggressive (since "+" is better than "0" as a payoff);
  3. If Israel chooses to be aggressive, Iran's best response is to be submissive (since "-" is better than "--" as a payoff); and
  4. If Israel chooses to be submissive, Iran's best response is to be aggressive (since "+" is better than "0" as a payoff).

In this scenario, there are no dominant strategies. Neither country has a strategy that is always better for them, no matter what the other country chooses to do. However, there are two Nash equilibriums (outcomes where both players are playing their best response), which occur when one country is aggressive, and the other is submissive.

The thing about the Iran-Israel hawk-dove game is that it isn't really a simultaneous game, as shown in the table above. It is a sequential game. Each player chooses whether to be aggressive or submissive, knowing what the other player chose to do previously. That sequential game is shown below. [*]

We can solve a sequential game using 'backward induction', which is essentially the same as the best response method, except we make sure we start with the last player, and work our way backwards through the game to work out what the first player should do. The resulting equilibrium that we find will be a 'subgame perfect Nash equilibrium'. In this case:

  1. If Israel chooses to be aggressive, Iran's best response is to be submissive (since "-" is better than "--" as a payoff) - now, since Iran would never choose to be aggressive when Israel has already been aggressive, Israel knows that the outcome will be that Iran is submissive;
  2. If Israel chooses to be submissive, Iran's best response is to be aggressive (since "+" is better than "0" as a payoff) - now, since Iran would never choose to be submissive when Israel has already been submissive, Israel knows that the outcome will be that Iran is aggressive;
  3. Israel will choose to be aggressive (since "+" is better than "-" as a payoff).

The subgame perfect Nash equilibrium is that Israel is aggressive, and Iran is submissive. As it turns out, that's sort of what happened. After an initial flurry of missile attacks, Iran stopped escalating the conflict.

One last thing to note is that this is a repeated game. Israel and Iran will find themselves in conflict often. The games as outlined above suggest that whichever country is the initial aggressor will end up getting their way, because the country moving second in the sequential game will be better off being submissive than retaliating. However, in repeated games, the outcome can often deviate from the equilibrium for strategic reasons.

Israel clearly doesn't want Iran to be aggressive, even though aggression would be good for Iran, if they moved first in this game. So, Israel wants to convince Iran not to make an aggressive first move. The only way that Israel can do that is to convince Iran that Iran would be worse off by making an aggressive first move. Israel needs to convince Iran that Israel will always retaliate with aggression. That would deter Iran from being aggressive as a first move. How does Israel achieve this? By developing a reputation for aggressively retaliating against any aggression. And indeed, that is what Israel has done (one need only look at Gaza or Lebanon for confirmation of this).

Israel's aggressive response to attacks by Iran, Gaza, and Lebanon is part of a strategic plan to deter future aggression against Israel. Many of us may not like it, but it's strategically rational. Whether it has a lasting effect remains to be seen.

*****

[*] I'm showing the game as having Israel move first. However, if you read Thomas's article, you'll see that 'who started it' is actually contested. I'm not taking a stand on that here, and in fact the game looks identical if Iran moves first.

Tuesday, 11 January 2022

Contract cheating, blackmail, and the end of game problem

When we designed the new ECONS101 paper some years ago, I tried desperately to include two weeks of game theory. Ultimately, it proved unworkable, but my rationale for trying to include it was that it would allow more time to explore repeated games. One aspect of that is the end of game problem. In a repeated game, the outcome may differ from the Nash equilibrium, because the players may realise that some other outcome benefits them more over many plays of the game. For example, in the repeated prisoners' dilemma, the dominant strategy equilibrium is for both prisoners to confess, but the best outcome is obtained if both prisoners remain silent (for an example of this, see this post on the drug dealers' dilemma). In a repeated game, players may be more likely to cooperate (and both prisoners stay silent) in the hopes of future cooperation. The American political scientist Robert Axelrod noted that players cooperate because of "the shadow of the future".

However, there is a difference between what happens when a game is infinitely repeated (it is played many times without end), and when a game is finitely repeated (played many times, but the players know when it will end). In a finitely repeated game, both players have an incentive to cheat in the last play of the game. Knowing that the other player will cheat in the last round, then both players realise that there is little point in cooperating in the second-to-last round of the game, so both players will cheat in that round. And knowing that, there is an incentive to cheat in the third-to-last round, and then the same logic applies to the fourth-to-last round, the fifth-to-last round, and so on all the way back to the first round of the repeated game. In a finitely repeated game, any cooperation breaks down, and neither player should be able to trust the other player to cooperate. Notice that this wouldn't be the case for an infinitely repeated game, because since neither player knows when the game will end, they never know when they should start cheating.

That brings me to the example of contract cheating, which involves a student outsourcing their assessment work to someone else to complete, usually for payment. Contract cheating has become a big issue at universities (see for example this 2017 article in The Conversation), and has been exacerbated by the shift to online teaching and assessment during the pandemic. Estimates suggest that around 8 percent of students may engage in contract cheating during their university studies. That might not sound like a lot, but that equates to hundreds of students at even a relatively small university like Waikato.

However, the up-front monetary payment that students face, and the chance that their cheating is detected and they are sanctioned, are not the only costs that students may face when they use contract cheating services. As this recent article in The Conversation notes, students may later be blackmailed by the contract cheating service, which threatens to reveal their cheating to the university unless further payments are made. In this new study by Jonathan Yorke, Lesley Sefcik, and Terisha Veeran-Colton (all Curtin University), published in the journal Studies in Higher Education (sorry, I don't see an ungated version online), 14 out of 587 students surveyed:

...stated that they directly or indirectly knew students who had been blackmailed by contract cheating services.

That's a disturbingly high incidence of blackmail, and something that students who engage these services should be concerned about. And if students were better advised of the risk of blackmail, contract cheating would probably decline. To see why, we can use some game theory and the end of game problem.

Consider a simple sequential game for a single assessment, as laid out in the decision tree (which we refer to as extensive form) below. The student makes a decision first, whether to engage in contract cheating or not. If they choose not to engage in contract cheating, the game ends, and their payoff (measured in utility) is based on the chances that they pass the assessment. If the student chooses to engage in contract cheating, then the cheating service chooses whether to blackmail the student or not. The payoff to the cheating service is profits from the fees or blackmail payment that the student pays.

There are three important aspects to this game. First, the Nash equilibrium of this game is for the student not to engage in contract cheating. To see why, we can use backward induction - essentially we start at the end of the game and work our way backwards, eliminating strategy choices that would not be chosen by each player. If the student chooses to engage in contract cheating, then the cheating service will choose to blackmail the student (because the payoff of 100 is greater than the payoff of 60). A rational student would then know that if they engage in contract cheating, they will get blackmailed, and their payoff will be -50. They won't get the payoff of 50, because the cheating service is going to blackmail them. Knowing this, the student is choosing between not engaging in contract cheating (and receiving a payoff of 15), and engaging in contract cheating (and receiving a payoff of -50). They are better off not engaging in contract cheating. That outcome is the subgame perfect Nash equilibrium in this game.

Second, many students are naïve. They don't realise that the cheating service can blackmail them. They think they are playing a different game, where they get to choose between the payoff of 50 (from engaging in contract cheating) and the payoff of 15 (from not engaging in contract cheating). The grey strategy and outcome are hidden from them. These naïve students think they are better off cheating, and will do so.

Third, this is actually a repeated game, because it is played over and over for each assessment that a student completes. That changes the incentives for the cheating service. Once a cheating service blackmails a student, the student probably isn't going to go back to that service. That means that the cheating service will receive a one-off payoff of 100. However, if the cheating service can encourage the student to return for more assessments, the cheating service will receive a payoff of 60 from every play of the game. Using the (made up) numbers from the game above, the cheating service only needs the student to pay twice in order to make them better off than they would be by blackmailing automatically. Strategically, the cheating service would therefore be better off overall by not blackmailing the student in each play of the game.

The repeated nature of the game affects naïve and rational students differently. The naïve students didn't realise that blackmail was an option at all, so from their perspective nothing has changed. The rational students may be tempted to engage in some contract cheating, even though they realise that the cheating service could engage in blackmail.

Of course, you can probably see where all this is going. The repeated game where the cheating service avoids engaging in blackmail only exists in an infinitely repeated game. However, this game is not infinitely repeated, because eventually the student is going to graduate, after which they won't need the contract cheating service any more. The contract cheating service has a strong incentive to string the student along, extracting fees from each assessment, and collecting evidence of the student's cheating, before blackmailing them in the last play of the game (which might even be after the student has graduated!).

In this finitely repeated game, the naïve students are hurt tremendously. At least if the game was not repeated, they get blackmailed in the first play. However, this finitely repeated game maximises the profits of the cheating service, by keeping the naïve student in the game until the very end. For the rational students, hopefully they are rational enough not only to recognise the possibility of blackmail, but also the end of game problem.

How can universities help students and mitigate the problem of blackmail from contract cheating services? Most universities now require students to complete an academic integrity module at the start of their studies, which impresses upon them that cheating (of various forms) is not allowed. One component of those programmes should certainly be (as discussed by Yorke et al.) to discuss with students the possibility of blackmail by contract cheating services, and how widespread (and growing) the practice is. At the very least, that will convert naïve students into more rational students by making them realise the full game that they are playing with the contract cheating services.

Tuesday, 16 February 2021

Combating cheating in online tests

From my perspective, the most challenging aspect of teaching during the pandemic lockdowns last year wasn't the teaching itself, it was dealing with students cheating in the online assessment. To give you some idea, I sent more students to the Student Discipline Committee in B Trimester 2020 than I had in the previous 10 years of teaching combined. All but one of those students ended up failing their paper. And I was not alone. The Student Disciplinary Committee faced a huge increase in workload, especially related to students using contract cheating websites to answer assessment questions for them.

Anyway, as you may expect, my experiences (and those of my colleagues) are not isolated examples. In a new paper in the Journal of Economic Behavior and Organization (ungated earlier version here), Eren Bilen (University of South Carolina) and Alexander Matros (Lancaster University) looked at cheating in online assessments. They use two examples to illustrate the pervasiveness of cheating: (1) students in an intermediate level class in Spring Semester 2020 (when lockdowns were introduced partway through the semester); and (2) online chess tournaments. They motivate their analysis with a simple game theoretic model, as shown below (the first payoff is to the student, and the second payoff is to the professor).


They note that in the sequential game:

It is easy to find a unique subgame perfect equilibrium outcome, where the student is honest and the professor does not report the student. Note that this is the best outcome for the professor and the second best outcome for the student.

To see why that is the subgame perfect Nash equilibrium, we can use backward induction. Essentially, we work out what the second player (the professor) will do first, and then use that to work out what the first player (the student) will do. In this case, if the student cheats, then we are moving down the left branch of the tree. The best option for the professor in that case is to report the student (since a payoff of 3 is better than a payoff of 2). So, the student knows that if they cheat, the professor will report them. Now, if the student doesn't cheat, then we are moving down the right branch of the tree. The best option for the professor in that case is not to report the student (since a payoff of 4 is better than a payoff of 1). So, the student knows that if they don't cheat, the professor will not report them. So, the choice for the student is to cheat and get reported (and receive a payoff of 1) or not cheat and not get reported (and receive a payoff of 3). Of course, the student will choose not to cheat. The subgame perfect Nash equilibrium here is that the student doesn't cheat, and the professor doesn't report them.

The problem with that analysis is that the professor doesn't know with certainty if the student has cheated or not. So, Bilen and Matros move onto a sequential game, as shown below. Even though the players make their choices sequentially, because the student's choice about whether to cheat or not is not revealed to the professor, it is as if the professor is making their choice about whether to report or not at the same time as the student. That makes this a simultaneous game.



Bilen and Matros note that, in this game:
This game has a unique mixed-strategy equilibrium, which means that the student and the professor should randomize between their two actions in equilibrium. Thus cheating as well as reporting is a part of the equilibrium.
To see why, we need to try to find the Nash equilibriums in this game, and to do that we can use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the textbook definition of Nash equilibrium). In this game, the best responses are:
  1. If the student chooses to cheat, the professor's best response is to report the student (since 3 is a better payoff than 2);
  2. If the student chooses not to cheat, the professor's best response is not to report the student (since 4 is a better payoff than 1);
  3. If the professor chooses to report the student, the student's best response is to not cheat (since 2 is a better payoff than 1); and
  4. If the professor chooses not to report the student, the student's best response is to cheat (since 4 is a better payoff than 3).
A Nash equilibrium occurs where both players' best responses coincide (normally I would track this with ticks and crosses, but since I didn't create the payoff table I haven't done so in this case). Notice that there isn't actually any case where both players are playing a best response. If the student cheats, the professor's best response is to report them. But if the professor is going to report the student, the student's best response is to not cheat. But if the student doesn't cheat, the professor's best response is not to report them. But if the professor doesn't report the student, the student's best response is to cheat. We are simply going around in a circle.

In cases such as this, we say that there is no Nash equilibrium in pure strategy. However, there will be a mixed strategy equilibrium, where the players randomise their choices of strategy. The student should cheat with some probability, and the professor should report the student with some probability. The optimal probabilities depend on the actual payoffs to the players (we could work it out in this case, but I'm not going to go that far today).

Anyway, Bilen and Matros conclude that in-person tests and examinations most closely resemble the sequential game, where the equilibrium is for students not to cheat, but that online tests most closely resemble the simultaneous game, where cheating will be much more common. And, their analysis supports that theory:
Using a simple way to detect cheating - timestamps from the students’ Access Logs - we identify cases where students were able to type in their answers under thirty seconds per question. We found that the solution keys for the exam were distributed online, and these students typed in the correct as well as incorrect answers using the solution keys they had at hand.

Then they present their proposed solution to the problem of online cheating:

In order to address this issue based on our theoretical models, we suggest that instructors present their students with two options: (1) If a student voluntarily agrees to use a camera to record themselves while taking an exam, this record can be used as evidence of innocence if the student is accused of cheating; (2) If the student refuses to use a camera due to privacy concerns, the instructor should be allowed to make the final decision on whether or not the student is guilty of cheating, with evidence of cheating remaining private to the instructor.

I'm not sure that I agree. The optimal solution would be one that returns to conditions where cheating is easy to detect, as in the sequential game above. A voluntary webcam doesn't do this, since students who want to cheat, as well as students who have privacy concerns, would opt out. The game would revert to the simultaneous game for those students, and some of those students would cheat.

My solution, which will be implemented if we go into lockdown and can't have a final examination in the trimester due to start in a couple of weeks, is to move to individual oral examinations. It is a lot harder to hide when you are put on the spot in a Zoom call and asked to answer a question chosen at random by the lecturer. Certainly, you can't use an online cheating website to provide the answer for you in such a situation! I haven't fully worked out the mechanics of how it would work (and hopefully I never have to implement it!), but it would seem to me to return the assessment to a sequential game.

On the other hand, the theory of asymmetric information and signalling does suggest that the webcam solution may work. The student knows whether they intend to cheat or not. The professor doesn't know. Whether the student is planning to cheat is private information. The student can reveal this information by the choices they make. Say that the professor offers a voluntary webcam policy, where students who agree set up a webcam such that they can be observed while they complete the test. Students who agree to the webcam are clearly those that weren't intending to cheat, because then they would surely be caught. In contrast, if a student was intending to cheat, they wouldn't agree to the policy. And so, the professor is able to reveal who the likely cheaters are, simply by offering the policy. The students signal whether they are cheaters by agreeing to have the webcam, or not.

That seems way too simple, and you could argue that it is somewhat coercive, forcing honest students to compromise their privacy to signal their honesty. It's going to be imperfect too, because it would be difficult to separate students who chose not to webcam from those that can't afford a webcam, or whose internet connection is too slow to support webcam use, and so on.

And invigilating the webcam solution is going to be extraordinarily expensive. You can't get AI to do the supervision (at least, not yet). And since Zoom can only support 25 faces on screen at a time, you would need at least one invigilator per 25 students. Probably you need more than that, because the invigilator needs to be able to see clearly what the student is doing (so unless they are using a 50" screen, you probably want a whole lot fewer than 25 images on-screen at a time). I think I'll stick to the oral examination solution.

Cheating in online assessments is clearly a serious problem. This is the main reason that I am highly sceptical about the move to online education. Until we can solve the serious academic integrity issues in a cost-effective way (and we currently don't have one), the quality of assessment (in terms of accurately ranking students or assessing their knowledge and skills) is so flow that it makes a mockery of the whole idea of teaching online.

Wednesday, 24 May 2017

Three reasons why tipping is a bad idea

Tipping has been in the news this week. Matt Heath started it with this article on Sunday, but then Deputy Prime Minister (and former waitress) Paula Bennett chimed in, saying "Overall I think the service in New Zealand is good, I always tip for excellent service and encourage others to too if we want standards to continue to improve" (at least, according to this article - I didn't read her letter to the Herald myself). Bennett's comments have stirred a lot of media interest (see here and here and here, for example). Now, as the voice of reason, I give you three reasons why tipping is a bad idea.

First, it's not rational if it's not already a social convention. To see why, we need to go through a little bit of game theory (which is good revision for my ECON100 students, since we did game theory in class last week). Consider a sequential game with two players: (1) the server, who can choose to give average service, or good service; and (2) the customer, who can choose to tip, or not, and makes their choice after the service decision of the server has already been revealed. Let's say that the basic outcome (average service and no tip) leads to a zero payoff for both players. Let's also assume that if the server gives good service, that increases the payoff to the customer by +6 (units of utility, or satisfaction), but comes at a cost to the server of -2 (units of utility). Finally, let's assume that if the customer chooses to tip, that reduces their payoff by 5, and increases the server's payoff by 5. The game is laid out in tree form (extensive form) below.


To find the subgame perfect Nash equilibrium here, we can use backward induction (similar to the best response method we use in a simultaneous game). Essentially, we work out what the second player (the customer) will do first, and then use that to work out what the first player (the server) will do. In this case, if the server gives good service, then we are moving down the left branch of the tree. The best option for the customer in that case is not to tip (since a payoff of +6 is better than a payoff of +1). So, the server knows that if they give good service, the customer is better off not tipping. Now, if the server gives average service, then we are moving down the right branch of the tree. The best option for the customer in that case is not to tip (since a payoff of 0 is better than a payoff of -5). So, the server knows that if they give average service, the customer is better off not tipping. Notice that the customer is better off not tipping no matter what the server does - not tipping is a dominant strategy for the customer. So, the choice for the server is to give good service (and receive a payoff of -2) or to give average service (and receive a payoff of 0). Of course, they will give average service. The subgame perfect Nash equilibrium here is that the server gives average service, and the customer doesn't leave a tip.

However, that analysis assumes that this is a non-repeated game. We know that if games are repeated, the outcome may be able to move away from the Nash equilibrium to an outcome that is better for all players (notice that the combination of good service and tipping is better for both players). How do we get to this alternative outcome? It relies on cooperation between the two players, and cooperation requires trust. The server has to trust that the customer will tip them, before they will agree to give good service. Can they trust the customer? Only if they have developed a relationship with that customer, and in most hospitality situations it is unlikely that a customer will encounter the same server again in the future (unless they are a regular). So, no trust. No cooperation. No tipping, and no good service.

Which brings me to social convention. One way to ensure cooperation from the customer is to make tipping a social convention, which has some social penalty attached to it. If it is frowned upon not to tip the server, to the extent that it becomes costly (in terms of moral costs or social costs, not financial costs) not to tip, then that changes the game. Say that the moral cost of not tipping is -6 units to the customer (since everyone who sees them not tipping the server then thinks the customer is a douchebag). This changes the game to this:


Now, where is the subgame perfect Nash equilibrium? If the server gives good service, the customer will tip (because +1 is better than 0). If the server gives average service, the customer will tip (because -5 is better than -6). Notice that tipping is now a dominant strategy for the customer. Knowing what the customer will do, the server will choose to give average service (since +5 is better than +3). The subgame perfect Nash equilibrium is now that the server gives average service, and the customer leaves a tip. We might wish that tipping would provide servers with an incentive to give good service, but that isn't always the case!

Nevertheless, tipping relies on a social convention, which is not the current convention in New Zealand. And developing new social conventions is not easy (although perhaps Paula Bennett is willing to give it a try in this case?).

The second reason why tipping is a bad idea is because of second-order effects. If customers have to tip the servers, this increases the cost of their meal. Since we know that demand curves are downward sloping, an increase in price will lead to lower quantity demanded - customers will demand fewer restaurant meals. If you doubt this point, then consider how many people you know (I'm sure there are at least some) who object to paying a surcharge for a meal on a public holiday, and so choose to either eat somewhere else (where there is no surcharge) or not to go out at all. Now, note that tipping is essentially the same as applying a surcharge to every restaurant meal.

Since the quantity of restaurant meals demanded will decrease, the number of servers required by restaurants also decreases. Tipping will make some servers better off (higher take-home pay), but will make others worse off (they no longer have a job). This has the same effect as raising the minimum wage, except the customers are paying the extra, rather than the employers. I'm not sure that's a trade-off that customers should be willing to accept.

The third reason why tipping is a bad idea is because it could be considered a form of corruption. If you doubt that, consider this example. Remember that the purpose of tipping is to reward the recipient for giving good service. Now, say that I'm pulled over by a police officer for driving through a stop sign, but the officer decides to let me off with a warning (seems unlikely, but let's run with it). The officer gave me good service - should I tip them?

The World Bank defines corruption as:
...the offering, giving, receiving or soliciting, directly or indirectly, anything of value to influence improperly the actions of another party.
Isn't tipping to reward good service providing something of value (money) to influence the actions of another party (to give you good service)? We could quibble over whether the influence is improper or not, I guess. But the general point is valid.

Anyway, now you have three reasons to use to explain why you shouldn't be tipping: (1) it's not rational (when there is no social convention for tipping); (2) it may make some servers worse off; and (3) it may be corrupt. You're welcome.

Friday, 19 May 2017

North Korea, nuclear missiles, and credible threats of extortion

Eric Rasmussen wrote an excellent post recently about North Korea's nuclear threat. I thought this would be topical to cover here, given that we discussed game theory in ECON100 this week and Rasmussen's post makes use of sequential games. Rasmussen writes:
Besides defense, though, the North Korean military does have another purpose: to make money...
The army can also make money by extortion.  North Korea’s army is too weak to engage in plunder by conquest, but it is strong enough to engage in demanding nuisance fees.   Would it be worth $20 billion per year to South Korea to avoid Seoul being bombarded? North Korea could be like the Barbary Pirates of 1800, who were enough of a nuisance to extract a goodly amount of their revenue as protection money, but so poor that that same amount was trivial to the European countries that paid it. The   United States,  having more principle and less monetary calculation, ironically, than the aristocratic Europeans, proved problematic on the shores of Tripoli and eventually France ended the game by conquering Algeria. However, the Barbary pirates did have a good run for their money.
The problem is making the threat to bombard Seoul credible.  The threat is credible if South Korea invades the North. If a South Korean invasion begins,  and  Kaesong is about to be occupied, North Korea has nothing to lose by destroying Seoul. If South Korea purposely bypasses Kaesong to avoid triggering that response and heads straight for Pyongyang, the Kim regime would see its demise and, again, might as well destroy Seoul and get a bit of revenge. Either way, the threat of bombardment is credible.
On the other hand, if North Korea simply says it will shell Seoul unless $20 billion is deposited in a certain Swiss bank account, that threat is not credible. If South Korea refuses, and North Korea shells Seoul, South Korea will respond by destroying the guns. Once the guns are gone,   South can conquer   North without fear of retaliation. North Korea will have almost literally “shot its wad.” South Korea may have lost 100,000 people, but that is small comfort for the Kim regime if it loses power. Thus, looking ahead, South can see that North will not retaliate and its $20 billion demand can be safely refused.
The game that Rasmussen describes is laid out in the figure below. The payoffs in the figure are (North Korea, South Korea). We can solve sequential games like this using backward induction - that is, we start with the last decision and work our way backwards. So, in this case the final decision is South Korea's - whether to Bomb Pyongyang, or not. If South Korea bombs Pyongyang, their payoff is -95, compared with -100 for not bombing Pyongyang. So, South Korea will bomb Pyongyang (because -95 is better than -100). Now, working backwards one step, we can work out whether North Korea will bombard Seoul. North Korea knows that South Korea will bomb Pyongyang if they bombard Seoul, so North Korea's payoffs are -200 if they bombard Seoul, or -1 if they don't. So, North Korea will choose not to bombard Seoul (because -1 is better than -200). Their threat to bombard Seoul if South Korea doesn't pay them $20 billion is not credible - North Korea wouldn't follow through on the threat.

Next, we can work out whether South Korea will choose to pay the $20 billion demand. If South Korea pays the $20 billion their payoff is -20, but if they don't pay their payoff is 1 (since North Korea will choose not to bombard Seoul). South Korea will choose not to pay the $20 billion (because 1 is better than -20). Finally, we can work out whether North Korea will threaten Seoul. If they issue the threat, their payoff is -1 (since South Korea will choose not to pay, and then North Korea will choose not to bombard Seoul), but if they don't issue the threat their payoff is 0. So, North Korea will not threaten Seoul (because 0 is better than -1). The subgame perfect Nash equilibrium is that North Korea doesn't threaten Seoul (and South Korea doesn't need to do anything in response, because the game ends right there). Note that Rasmussen has tracked all of the best responses as arrows in the figure.

Rasmussen then goes on to describe how the game would change if North Korea develops a nuclear arsenal:
...nukes are good for extortion in themselves and a good backup for artillery. Imagine now that Kim has nuclear missiles pointed at Seoul. This changes the game... There is now a new move at the end, Nuke Seoul or Not. Many of the arrows change, though, because that last move changes everything.
The game with nuclear weapons is in the figure below. Again, we can solve this game with backward induction. Now the final decision in the game is North Korea's - whether to Nuke Seoul, or not. If North Korea nukes Seoul, their payoff is -180, compared with -200 for not nuking Seoul. So, North Korea will nuke Seoul (because -180 is better than -200). Now, working backwards one step, we can work out whether South Korea will bomb Pyongyang. This time, if South Korea bombs Pyongyang their payoff is -195 (since North Korea will nuke Seoul), but their payoff is -100 if they don't bomb Pyongyang. So, South Korea will not bomb Pyongyang (because -100 is better than -195). Next, we can work out whether North Korea will bombard Seoul. North Korea knows that South Korea will not bomb Pyongyang if they bombard Seoul, so North Korea's payoffs are 2 if they bombard Seoul, or -1 if they don't. So, North Korea will now choose to bombard Seoul (because 2 is better than -1). If North Korea has nuclear weapons, notice that their threat to bombard Seoul is now credible - they will follow through on it.


Next, we can work out whether South Korea will choose to pay the $20 billion demand. If South Korea pay the $20 billion their payoff is -20, but if they don't pay their payoff is now -100 (since North Korea will choose to bombard Seoul, and then South Korea will choose not to bomb Pyongyang because Seoul would then get nuked). South Korea will choose to pay the $20 billion (because -20 is better than -100). Finally, we can work out whether North Korea will threaten Seoul. If they issue the threat, their payoff is 20 (since South Korea will choose to pay them), but if they don't issue the threat their payoff is 0. So, North Korea will now threaten Seoul (because 20 is better than 0). The subgame perfect Nash equilibrium is that North Korea threatens Seoul, and South Korea pays the $20 billion.

So, that provides one more reason (if any were needed) why South Korea (and its allies) should be working hard to prevent North Korea from developing nuclear weapons. Because North Korea could then use them for extortion.


Tuesday, 31 May 2016

The toilet seat game

When I blogged my review of William Nicolson's "The Romantic Economist - A Story of Love and Market Forces" last year, I mentioned that I would definitely use the toilet seat game as an example in ECON100 this year. However, the game theory topic is already pretty full, so I didn't manage to squeeze it in. So instead, I thought I would talk about it here.

First, a little bit of background. William tells us about one of his relationships, with Sarah, which hit a bit of a rocky patch with respect to toilet seats, and whether they should be left up or down. You may laugh, but there's been several papers that have explored the game theoretical implications of the toilet seat game (see here and here for two examples).

In the case of Will and Sarah, the decision-making can be thought of as a sequential game, with three decision-making nodes. In the first node of the game, Will decides whether to leave the seat up or down. If he leaves the seat down, the game ends (because Sarah is happy). But if he leaves the seat up, Sarah gets angry. She then has the choice of whether to forgive Will, or punish him by nagging or throwing a minor tantrum. If Sarah forgives Will, then the game ends (and Will is pretty happy, because he got to leave the seat up and avoid the nagging). If instead Sarah punishes Will, then Will has the choice of agreeing to change his behaviour for the future, or ending the relationship. The game as a decision tree (extensive form) is presented below. The payoffs to Will and Sarah are simply ranked from 1 (best) to 4 (worst). Note that in this version of the game, ending the relationship is a worse outcome for Sarah than for Will. [*]

We can solve for the (subgame perfect) Nash equilibrium by using backward induction - essentially we start at the end of the game and work our way backwards, eliminating strategy choices that would not be chosen by each player. In this case, the last play is by Will. He can choose to end the relationship (and receive his 3rd best payoff), or change his behaviour (and receive his worst payoff). So, clearly he would choose to end the relationship. Now, working our way back to Sarah's choice, if she punishes Will then we know that Will is going to end the relationship. So Sarah can choose to forgive Will (and receive her 3rd best payoff), or punish him (which results in Will ending the relationship and Sarah receiving her worst payoff). Given that choice, Sarah will choose to forgive Will. Working our way back to the first node then, Will can choose to leave the seat down (and receive his 2nd best payoff) or leave the seat up (knowing that Sarah will forgive him, and he will receive his best payoff). Given that choice, Will is going to leave the seat up. So, the subgame perfect Nash equilibrium in this case is that Will leaves the seat up, and Sarah forgives him.

Will then goes on to point out that the outcome of this game crucially depends on how Will feels about Sarah. If he really wants to be with Sarah, then the game changes. Now say that ending the relationship is the worst outcome for both Sarah and Will, as showing in the decision tree below.


Solving for equilibrium in this case (again using backward induction), we find that Will would agree to change his behaviour (because that provides him with his 3rd best payoff, compared with the worst payoff if he instead chose to end the relationship). Now, knowing that Will is going to change his behaviour, Sarah would choose to punish him if he leaves the seat up (because she would receive her second best payoff when he agrees to change, which is better than her third best payoff, which she would receive by forgiving Will). Now, knowing that Sarah will punish him, Will decides to leave the seat down (and receive his second best payoff, which is better than the third best payoff that he would receive if he left the seat up, because Sarah would then punish him and he would then change his behaviour).

All in all, this is a great example of a sequential game in action. Of course, the game as presented above is necessarily simplified - in fact this game is a repeated game, so the Nash equilibrium is more complex (as shown in the two papers linked at the beginning). But a great example nonetheless.

*****

[*] Careful readers of Nicolson's book will notice that I have altered the payoffs in these games from those that appear in the book. Specifically, the payoff for Sarah in the case of the Up/Punish/Change outcome is that this is Sarah's second best payoff (whereas in the book it is her third, which is equal to the Up/Forgive outcome). I expect this was a slight error in the book.