Showing posts with label Chicken game. Show all posts
Showing posts with label Chicken game. Show all posts

Thursday, 13 March 2025

Hawks, doves, Israel and Iran

In The Conversation last October, Andrew Thomas (Deakin University) discussed the recent (at that time) military flare-up between Iran and Israel, likening it to a 'game of chicken':

Israel’s strike on military targets in Iran over the weekend is becoming a more routine occurrence in the decades-long rivalry between the two states...

There is a reason why direct military strikes between nations are rare, even between sworn enemies. When attacking another state, it is difficult to know exactly how they will respond, though a retaliatory strike is almost often expected.

This is because defence forces are not just used for fighting and winning wars – they are also vital to deterring them. When a fighting force is attacked, it’s important for it to strike back to maintain the perception it can deter future attacks and make a display of its capabilities. This is what is happening right now between Israel and Iran – neither side wants to appear weak.

If this is the case, where does the escalation end? De-escalation is essentially a game of chicken – one side has to be content with not responding to an attack to take the temperature down.

My ECONS101 class has been covering game theory this week, including the chicken game. In the traditional game of chicken there are two rivals in cars, one at each end of the same street. They drive towards each other at top speed, and whichever rival swerves away first loses the game. So, each rival can choose to speed ahead or swerve away, and each would prefer to speed ahead and win the game. However, the problem is that if both simply keep speeding ahead, it will end in a disastrous crash.

I also recently read the book Hidden Games, by Moshe Hoffman and Erez Yoeli (which I reviewed here). Hoffman and Yoeli have an interesting section in the book on the hawk-dove game, which essentially the chicken game but with a slightly different motivating context. In the hawk-dove game, two rivals are competing over some resource. Each rival can choose to be aggressive or submissive, and whichever rival is more aggressive will win the resource. Each rival would prefer to be aggressive and win the resource. However, if both are aggressive, it ends in a massively disastrous battle.

Coming back to the case of Iran and Israel, this is clearly an example of the hawk-dove game (or the chicken game, if you prefer). This game is laid out in the payoff table below, where the strategies for Israel and Iran are to be aggressive, or submissive. The payoffs are expressed as "+" for good outcomes, and "-" for bad outcomes (and "--" is particularly bad), while zero is a neutral payoff.

To find the Nash equilibrium in this game, we use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium). In this game, the best responses are:

  1. If Iran chooses to be aggressive, Israel's best response is to be submissive (since "-" is better than "--" as a payoff - in other words, taking a bit of punishment is better than a massively disastrous war) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If Iran chooses to be submissive, Israel's best response is to be aggressive (since "+" is better than "0" as a payoff);
  3. If Israel chooses to be aggressive, Iran's best response is to be submissive (since "-" is better than "--" as a payoff); and
  4. If Israel chooses to be submissive, Iran's best response is to be aggressive (since "+" is better than "0" as a payoff).

In this scenario, there are no dominant strategies. Neither country has a strategy that is always better for them, no matter what the other country chooses to do. However, there are two Nash equilibriums (outcomes where both players are playing their best response), which occur when one country is aggressive, and the other is submissive.

The thing about the Iran-Israel hawk-dove game is that it isn't really a simultaneous game, as shown in the table above. It is a sequential game. Each player chooses whether to be aggressive or submissive, knowing what the other player chose to do previously. That sequential game is shown below. [*]

We can solve a sequential game using 'backward induction', which is essentially the same as the best response method, except we make sure we start with the last player, and work our way backwards through the game to work out what the first player should do. The resulting equilibrium that we find will be a 'subgame perfect Nash equilibrium'. In this case:

  1. If Israel chooses to be aggressive, Iran's best response is to be submissive (since "-" is better than "--" as a payoff) - now, since Iran would never choose to be aggressive when Israel has already been aggressive, Israel knows that the outcome will be that Iran is submissive;
  2. If Israel chooses to be submissive, Iran's best response is to be aggressive (since "+" is better than "0" as a payoff) - now, since Iran would never choose to be submissive when Israel has already been submissive, Israel knows that the outcome will be that Iran is aggressive;
  3. Israel will choose to be aggressive (since "+" is better than "-" as a payoff).

The subgame perfect Nash equilibrium is that Israel is aggressive, and Iran is submissive. As it turns out, that's sort of what happened. After an initial flurry of missile attacks, Iran stopped escalating the conflict.

One last thing to note is that this is a repeated game. Israel and Iran will find themselves in conflict often. The games as outlined above suggest that whichever country is the initial aggressor will end up getting their way, because the country moving second in the sequential game will be better off being submissive than retaliating. However, in repeated games, the outcome can often deviate from the equilibrium for strategic reasons.

Israel clearly doesn't want Iran to be aggressive, even though aggression would be good for Iran, if they moved first in this game. So, Israel wants to convince Iran not to make an aggressive first move. The only way that Israel can do that is to convince Iran that Iran would be worse off by making an aggressive first move. Israel needs to convince Iran that Israel will always retaliate with aggression. That would deter Iran from being aggressive as a first move. How does Israel achieve this? By developing a reputation for aggressively retaliating against any aggression. And indeed, that is what Israel has done (one need only look at Gaza or Lebanon for confirmation of this).

Israel's aggressive response to attacks by Iran, Gaza, and Lebanon is part of a strategic plan to deter future aggression against Israel. Many of us may not like it, but it's strategically rational. Whether it has a lasting effect remains to be seen.

*****

[*] I'm showing the game as having Israel move first. However, if you read Thomas's article, you'll see that 'who started it' is actually contested. I'm not taking a stand on that here, and in fact the game looks identical if Iran moves first.

Sunday, 24 March 2024

In The Three Body Problem trilogy, Wallfacer Rey Diaz needed to better understand game theory

Regular readers of this blog may have noticed that I haven't posted a book review in a while. That's because I've been reading Cixin Liu's Three Body Problem trilogy (technically, the Remembrance of Earth's Past trilogy, an adaptation of which has just been released on Netflix as The Three Body Problem). I'm currently reading the second book, The Dark Forest.

Warning: Spoiler alert!

To give you some context, in the first book of the trilogy, Earth made contact with an alien civilisation, the Trisolarans. The Trisolaran fleet is currently on its way to Earth, in order to conquer us. Their homeworld is about to be destroyed, and their only hope of survival is to take over another planet. The fleet will take some 400 years to arrive, so Earth has some time to prepare. However, the Trisolarans have advanced technology, including deploying sophons, which are able to prevent Earth from conducting basic research in physics and other areas. So, Earth is stuck in a low-technology state, awaiting the arrival of the Trisolaran fleet. Even worse, the sophons can watch anything that happens on Earth and relay the information back to the Trisolarans, so Earth's preparations will be known to the Trisolaran fleet. To combat this, in the second book, Earth appoints four 'Wallfacers', who are given access to almost unlimited resources to execute plans that are known only to themselves, hidden from the rest of the Earth's population (and to the sophons, because the sophons can't read minds).

The second book of the trilogy is devoted to the Wallfacers and their plans (admittedly, I haven't finished reading it yet). I want to focus on the plans of Wallfacer Rey Diaz, whose plan involved planting large solar hydrogen bombs on Mercury, which when detonated would set off a chain reaction, destroying most of the solar system, including Earth. Diaz's plan was to negotiate with the Trisolaran fleet, warning them that if they didn't divert, Earth would be destroyed, sealing the fates of both the human and Trisolaran populations.

However, Wallfacer Rey Diaz's strategy is flawed. He needs to understand some basic game theory. To see why, consider the game shown in the payoff table below. The two players are Earth and the Trisolaran fleet (we'll assume that Diaz would choose strategy on behalf of Earth). Earth's two strategies are to blow up Mercury (detonate) or not. The Trisolaran fleet's two strategies are to continue to Earth, or divert. If Earth blows up Mercury, then Earth becomes extinct, regardless of what the Trisolaran fleet does. If Earth doesn't detonate Mercury, then Earth loses if the Trisolaran fleet continues, and wins if the Trisolaran fleet diverts. If the Trisolaran fleet diverts, they become extinct. If they continue to Earth, they become extinct if Earth blows up Mercury, but win if Earth does not.

To find the Nash equilibrium in this game, we use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium). In this game, the best responses are:

  1. If the Trisolaran fleet continues to Earth, Earth's best response is to not detonate (since losing is a better payoff than extinction [*]) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If the Trisolaran fleet diverts, Earth's best response is to not detonate (since winning is a better payoff than extinction);
  3. If Earth chooses to detonate, the Trisolaran fleet's best response is either option (since both payoffs are the same - extinction - both are best responses); and
  4. If Earth chooses not to detonate, the Trisolaran fleet's best response is to continue to Earth (since winning is a better payoff than extinction).

Note that Earth's best response is always to choose not to detonate. This is their dominant strategy. A player would always choose to play their dominant strategy, because choosing the other strategy makes them unambiguously worse off. And the Trisolarans would know this. This is what Wallfacer Rey Diaz gets wrong in his strategy. Earth won't blow Mercury up, and the Trisolarans know this, so there is no leverage for Earth in the negotiations.

The Trisolaran fleet has a weakly dominant strategy. Notice that continuing to Earth is always the Trisolaran fleet's best response. However, diverting is a best response if Earth chooses to detonate. So, continuing to Earth is not always better for the Trisolaran fleet, but it is never worse than the other strategy.

The single Nash equilibrium occurs where both players are playing a best response (where there are two ticks), which is where all Earth chooses not to detonate, and the Trisolaran fleet continues to Earth. It is little wonder then, that when Rey Diaz's strategy was revealed, the Earth governments were not happy. Not only was his strategy imperilling the Earth to the same extent as the Trisolarans, it was a strategy that simple game theory shows would not have succeeded.

*****

[*] You may wonder what the difference between losing and extinction is. Earth could lose, but some humans remain alive as slaves, or otherwise escape the planet before the Trisolarans arrive. It's not a great outcome, but better than extinction.

Tuesday, 1 August 2023

The Tour de France, public goods, and the chicken game

I finally finished watching this year's Tour de France on Sunday. Yes, I was a week behind. That's because I was overseas when it started, and it took me that long to catch up (with big thanks to Sky On Demand!). Jonas Vingegaard well deserved his win. The individual time trial he rode on Stage 16 was amazing to watch (even if his team Jumbo Visma says so themselves).

Anyway, this is a blog about economics. Sports provide lots of great examples of economics in action, because economics is ultimately about choices, and so are sports. One striking example of economics in action in cycling road races occurs when there is a breakaway, and it is getting close to the finish line. The riders in the breakaway face a difficult choice. They can ride hard at the front of the breakaway, ensuring that the breakaway won't be caught by the peloton, and one of the breakaway riders will surely win the race. Or they can hold back, riding in the slipstream of the rider who is riding at the front, which lets them conserve energy for a sprint finish, but at the risk that the peloton catches them.

This exact scenario played out in Stage 18 of the Tour de France this year, with three riders approaching the finish. Victor Campenaerts rode hard towards the finish, ensuring the breakaway would succeed. However, it was Kasper Asgreen who won the stage, having conserved his energy for the final sprint among the breakaway riders.

Let's think about the incentives for a breakaway rider. Riding hard is a public good. It is non-rival (one cyclist benefiting from a rider riding hard at the front of the breakaway doesn't reduce the amount of the benefit available for the other riders in the breakaway) and non-excludable (if a rider is riding hard at the front of the breakaway, they can't easily prevent the other breakaway riders from sitting in their slipstream and conserving their energy).

Public goods, like riding hard at the front of the breakaway group, suffer from a free rider problem (pun intended!). Other riders can benefit from the front rider's hard work, without paying any of the cost themselves. It is difficult for a rider to justify riding hard at the front if other riders are unwilling to contribute, since they face all of the cost of riding hard, but the benefit (in terms of a better chance of winning the race) goes to the other riders (the free riders).

Ordinarily, the provision of public goods breaks down. They cannot be privately provided, because of the free rider problem. In this case though, cycling has developed norms that ensure some cooperation within the breakaway group. The riders tend to take turns at the front of the breakaway group, helping to increase the chances of success. However, the closer the race gets to the finish, the greater the incentives to free ride become. Regular cycling fans will no doubt remember many instances where a breakaway group has been caught, within sight of the finish line, because they failed to work together.

Another way of thinking about the incentives within a breakaway group is to use game theory. To make the problem simpler, let's say that the breakaway group only consists of two riders, and there are two strategies: (1) to ride hard; or (2) to hold back. We'll assume each rider makes their decision just once, and they make their decisions at the same time (a simultaneous game).  The payoffs for this scenario are shown in the table below. If both riders ride hard, they have a 50% chance of winning the race (since they will both be equally tired). If one rider rides hard and the other holds back, the rider that holds back wins the race for sure. If both riders hold back, then they are caught by the peloton, and neither of them wins (and they don't even finish in the top two in the race). What will happen?

To find the Nash equilibrium in this game, we use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium). In this game, the best responses are:

  1. If Rider B chooses to ride hard, Rider A's best response is to hold back (since winning for sure is better than a 50/50 chance of winning) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If Rider B chooses to hold back, Rider A's best response is to ride hard (since losing and finishing in the top two is better being caught by the peloton and finishing much lower in the order);
  3. If Rider A chooses to ride hard, Rider B's best response is to hold back (since winning for sure is better than a 50/50 chance of winning); and
  4. If Rider A chooses to hold back, Rider B's best response is to ride hard (since losing and finishing in the top two is better being caught by the peloton and finishing much lower in the order).

In this scenario, there are no dominant strategies. Neither rider has a strategy that is always better for them, no matter what the other rider chooses to do. However, there are two Nash equilibriums (outcomes where both players are playing their best response), which occur when one rider rides hard, and the other holds back. Neither rider will want to be the rider that rides hard, so both may be holding out hoping that the other rider will ride hard. This is the free rider problem described earlier. This game is an example of the chicken game (which I have discussed here). If both riders hold back, hoping that the other rider will ride hard, both riders will be caught by the peloton.

The chicken game is an example of a coordination game. To end up at one of the equilibriums (or another), the players need to coordinate their actions. However, in this case neither rider really wants to coordinate on the other rider's preferred equilibrium. Both really want to hold back, especially closer to the finish line, which is why the breakaway can often be caught.

Riders are motivated by the chance to win the race. That is why breakaway groups form in the first place. However, the incentives outlined above work against the breakaway succeeding. And riders are aware of these issues. One thing that often happens is that, towards the end of a race, one rider will ride especially hard, breaking away from the breakaway group. There is no free rider problem when a rider is riding by themselves. Sadly, solo breakaways are seldom successful (except on mountain stages), because the effort required to remain clear from a group of breakaway riders who suddenly become more motivated to work together and catch the solo breakaway rider is very high. The solo breakaway rider is often caught, after which the chicken game and free riding begins again.

One thing that can increase the success of a breakaway is to have multiple teammates in the breakaway group. Teammates are more likely (but not certain) to be able to coordinate their strategies, and work together, reducing the free riding problem. That's why riders in the peloton are more vigilant and energetic in chasing down an early breakaway group that has multiple riders from the same team. Most of the time, a breakaway group will only go clear if every rider in the group is from a different team. Riders in the peloton don't want the breakaway to succeed, and having all breakaway riders from different teams decreases the chance that a rider from the breakaway wins the race.

There is a lot of strategy in sports, and cycling is no exception. There are also a lot of choices for athletes to make, and choices involves trade-offs. That, along with the transparent rules and the obvious goals of the athletes involved (they want to win), is why sports can provide a lot of useful illustrations of economic concepts.

Monday, 10 April 2023

The game theory of an AI pause

My news feed has been dominated over the last week by arguments both for and against a pause on AI development, prompted by this open letter by the Future of Life Institute., which called for:

AI labs to immediately pause for at least 6 months the training of AI systems more powerful than GPT-4.

Chas Hasbrouck has an excellent post summarising the various views on AI held by different groups (and examples of the people belonging to each group). Tyler Cowen then suggested on the Marginal Revolution blog that we should be considering the game theory of this situation (see also his column on Bloomberg - as Hasbrouck notes, Cowen is one of the 'Pragmatists'). I want to follow up Cowen's suggestion, and look at the game theory. However, things are complicated a little, because it isn't clear what the payoffs are in this game. There is so much uncertainty. So, in this post, I present three different scenarios, and work through the game theory of each of them. For simplicity, each game has two players (call them Country A and Country B), and each player has two strategies (pause development on AI, or speed ahead).

Scenario #1: AI Doom with any development

In this scenario, if either country speeds ahead and the other doesn't, the outcomes are bad, but if both countries speed ahead, the planet faces an extinction-level event (for humans, at the least). The payoffs for this scenario are shown in the table below.

To find the Nash equilibrium in this game, we use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium). In this game, the best responses are:

  1. If Country B chooses to pause development, Country A's best response is to pause development (since a payoff of 0 is better than a payoff of -5) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If Country B chooses to speed ahead, Country A's best response is to pause development (since a payoff of -10 is better than extinction);
  3. If Country A chooses to pause development, Country B's best response is to pause development (since a payoff of 0 is better than a payoff of -5); and
  4. If Country A chooses to speed ahead, Country B's best response is to pause development (since a payoff of -10 is better than extinction).

In this scenario, both countries have a dominant strategy to pause development. Pausing development is always better for a country, no matter what the other country decides to do (pausing development is always the best response).

For anyone who believes in this scenario, pausing development will seem like a no-brainer, since it is a dominant strategy.

Scenario #2: AI Doom if everyone speeds ahead

In this scenario, if both countries speed ahead, the planet faces an extinction-level event (for humans, at the least). However, if only one country speeds ahead, then AI alignment can keep up, preventing the extinction-level event. The country that speeds ahead earns a big advantage. The payoffs for this scenario are shown in the table below.

Again, let's find the Nash equilibrium using the best response method. In this game, the best responses are:

  1. If Country B chooses to pause development, Country A's best response is to speed ahead (since a payoff of 10 is better than a payoff of 0);
  2. If Country B chooses to speed ahead, Country A's best response is to pause development (since a payoff of -2 is better than extinction);
  3. If Country A chooses to pause development, Country B's best response is to speed ahead (since a payoff of 10 is better than a payoff of 0); and
  4. If Country A chooses to speed ahead, Country B's best response is to pause development (since a payoff of -2 is better than extinction).

In this scenario, there is no dominant strategy. However, there are two Nash equilibriums, which occur when one country speeds ahead, and the other pauses development. Neither country will want to be the country that pauses, so both will be holding out hoping that the other country will pause. This is an example of the chicken game (which I have discussed here). If both countries speed ahead, hoping that the other country will pause, we will end up with an extinction-level event.

For anyone who believes in this scenario, pausing development will seem like a good option, even if only one country will pause development. However, no country is going to want to willingly buy into pausing development.

Scenario #3: AI Utopia

In this scenario, if both countries speed ahead, the planet reaches an AI utopia. The fears of an extinction-level event do not play out, and everyone is gloriously happy. However, if only one country speeds ahead, then the outcomes are good, but not as good as they would be if both countries sped ahead. Also, the country that speeds ahead earns a big advantage. The payoffs for this scenario are shown in the table below.

Again, let's find the Nash equilibrium using the best response method. In this game, the best responses are:

  1. If Country B chooses to pause development, Country A's best response is to speed ahead (since a payoff of 10 is better than a payoff of 0);
  2. If Country B chooses to speed ahead, Country A's best response is to speed ahead (since utopia is better than a payoff of -2);
  3. If Country A chooses to pause development, Country B's best response is to speed ahead (since a payoff of 10 is better than a payoff of 0); and
  4. If Country A chooses to speed ahead, Country B's best response is to speed ahead (since utopia is better than a payoff of -2).

In this scenario, both countries have a dominant strategy to speed ahead. Speeding ahead is always better for a country, no matter what the other country decides to do (speeding ahead is always the best response).

For anyone who believes in this scenario, speeding ahead will seem like a no-brainer, since it is a dominant strategy.

Which is the 'true' scenario? I have no idea. No one has any idea. We could ask ChatGPT, but I strongly suspect that ChatGPT will have no idea as well. [*] What the experts believe we should do depends on which of the scenarios they believe is likely to be playing out. Or perhaps, with a chance that any of the three scenarios (or any other of millions of other potential scenarios with different players and payoffs) is playing out, perhaps the precautionary principle should apply? The problem there, though, is if any country pauses development, the best response in any of the scenarios except the first one is for other countries to speed ahead. So, unless all countries can be convinced to apply the precautionary principle, pausing development is simply unlikely.

We live in interesting times.

*****

[*] Actually, I tried this, and ChatGPT refused to offer an opinion, instead it said: "...it is crucial that policymakers and stakeholders work together to develop standards and guidelines for responsible AI development and deployment to minimize potential risks and maximize benefits for society as a whole." Thanks ChatGPT.

Tuesday, 21 March 2023

Spotify has artists playing chicken, but they can fight back if they can cooperate

The New Zealand Herald reported today:

Spotify has recently faced backlash over its newly-implemented Discovery Mode program.

The initiative, which gives artists greater exposure on the platform in exchange for a lower royalty rate, was announced during the company’s Stream On event in March 2021 and has continued to be criticised all the way up to its 2023 launch.

Under Discovery Mode, artists or their teams can submit tracks for consideration to be included on Spotify’s radio and autoplay features. In exchange for this greater algorithmic exposure, they agree to receive a lower royalty rate for streams of their music.

For some, it’s an inventive new way to link potential fans to new music, but others in the industry believe it to be yet another way to shave the pay cheque of hardworking musicians whose work is the lifeblood of an app that rakes in billions of dollars each year.

This is a smart ploy by Spotify. As noted by DJ Luca Lush on Twitter:

Ideally for spotify, EVERYONE opts in, they take 30% more revenue & no one gets more plays

It's not clear that every artist would opt into Discovery Mode though. To see why, consider the decision as part of a simultaneous game, played by some artist (Artist A) and all other artists. The game is laid out in the payoff table below, with the payoffs measured as a percentage of the 'normal' level of royalties. If all artists (including Artist A) choose no Discovery Mode, then they all continue to receive the normal level of royalties. If all artists (including Artist A) choose Discovery Mode, then they all lose 30 percent of their income (as Luca Lush noted). However, if Artist A chooses Discovery Mode and all others do not, Artist A benefits greatly (let's say that their royalties go up by 50 percent - in the New Zealand Herald article, Spotify says that "artists have seen an average 50 per cent increase in saves" when participating in Discovery Mode), and other artists are negatively affected, but only slightly (because even when Artist A's Spotify streams increase a lot, that doesn't much reduce every other artist's streams). On the other hand, if Artist A chooses not to participate in Discovery Mode and all other artists do, it is the other artists that benefit greatly, and Artist A is made worse off. [*]

To find the Nash equilibrium in this game, we use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium). In this game, the best responses are:

  1. If Artist A chooses not to participate in Discovery Mode, the other artists' best response is to participate in Discovery Mode (since a payoff of 150 is better than a payoff of 100) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If Artist A chooses to participate in Discovery Mode, the other artists' best response is not to participate in Discovery Mode (since a payoff of 99 is better than a payoff of 70);
  3. If the other artists choose not to participate in Discovery Mode, Artist A's best response is to participate in Discovery Mode (since a payoff of 150 is better than a payoff of 100); and
  4. If the other artists choose to participate in Discovery Mode, Artist A's best response is not to participate in Discovery Mode (since a payoff of 95 is better than a payoff of 70).

Notice that there are two Nash equilibriums in this game - where Artist A chooses to participate in Discovery Mode, and every other artist does not, and where Artist A chooses not to participate in Discovery Mode, and every other artist does.

We could repeat this exercise for any number of additional artists, rather than Artist A. We would come out with the same outcome. We could even try this as a multi-player game. We would find something similar. The artists are all better off if they can participate in Discovery Mode, but not too many of the other artists do so. Every artist would want to be participating in Discovery Mode. However, if all (or a large proportion) of them choose to participate in Discovery Mode, then every participant is made worse off. This is an example of the game of chicken. Two drivers driving towards each other can choose to speed ahead, or swerve out of the way. Both prefer to speed ahead, because they are trying to win the game of chicken. However, if they both follow that strategy, it all ends in a fiery crash (for a more complete explanation, see this post).

Usually, in the game of chicken, a player can get the outcome that they prefer if they can make a credible commitment to their strategy. In the classic game of chicken, a driver could commit to speeding ahead by disabling their brakes and throwing the steering wheel out the window. That is a pretty showy way of demonstrating that the driver won't change their strategy of speeding ahead.

However, in the game that Spotify has set up, there are too many players for a credible commitment to scare others off. Most artists will instead be thinking, 'I'm sure that one more artist choosing Discovery Mode isn't going to be the one that destroys the payoffs for everyone, so why shouldn't I?'. That's the sort of thinking that ends in a fiery crash, with all artists earning 30 percent lower royalties.

A cynic would interpret this as Spotify's strategy all along. They are profit maximising by steering the artists into a game of chicken that leads to all participating artists receiving lower royalties. However, the artists can fight back. This is a repeated game. In a repeated game, the players can cooperate in order to obtain a better outcome for them all. By cooperating, and choosing not to participate in Discovery Mode, the artists would be made better off collectively. This is essentially what the artists who have spoken out against Discovery Mode are trying to do. They are trying to coordinate a cooperative response that sees all artists boycotting Discovery Mode, which would be for the betterment of them all.

This sort of cooperative outcome is only possible if the artists can trust each other. There is an incentive for any artist to cheat on the agreement. That's because, if Artist A (or any other artist) knows for sure that the other artists will not participate in Discovery Mode, then Artist A can participate and make themselves better off. Once that starts to happen, the cooperative agreement can quickly break down.

The questions now are, will the artists be able to agree not to participate, and if they do, will they be able to maintain trust and cooperation?

*****

Another way of thinking about the payoffs in this game is to recognise that Artist A is probably made wildly worse off by every other artist opting into Discovery Mode. The game with these new payoffs is shown below.

The best responses for the other artists are unchanged. However, for Artist A, the best responses are now:

  1. If the other artists choose not to participate in Discovery Mode, Artist A's best response is to participate in Discovery Mode (since a payoff of 150 is better than a payoff of 100);
  2. If the other artists choose to participate in Discovery Mode, Artist A's best response is to participate in Discovery Mode (since a payoff of 70 is better than a payoff of 25).

Notice that now, Artist A has a dominant strategy to participate in Discovery Mode. Participating in Discovery Mode is better for Artist A, no matter what the other artists do. They should always choose to participate in Discovery Mode. And this would apply to any other artist, if we replaced Artist A with them instead. This provides an even stronger incentive for artists to participate in Discovery Mode than in the chicken game shown earlier (it wouldn't be a chicken game any more, but much more like a prisoners' dilemma game with multiple players). However, the other points I make about the repeated game, cooperation and trust, all still apply to this version of the game as well.

Tuesday, 28 April 2020

CEOs playing games

Experimental economics is incredibly useful, because it allows economists to study decision-making in circumstances when basically all of the key parameters to the decision are controlled. However, one of the main problems with experimental economics is that the study population is often made up of students (see here and here for previous posts on this topic). So, it's particularly interesting when experimental economics makes us of samples made up of 'real people'.

For instance, in a new article (open access) published in the journal Experimental Economics, HÃ¥kan Holm (Lund University), Victor Nee (Cornell University), and Sonja Opper (Lund University), report on an experiment they conducted with Chinese CEOs and "comparable people in professional roles". Their sample size is quite large for this type of study - they have 200 CEOs and 200 other professionals.

Their experiment involves game theory - essentially, the research participants were asked to choose actions in three games: (1) the prisoners' dilemma (which I have written about before, most recently here); (2) a 'battle of the sexes' coordination game (I have written about coordination games too, see for example here); and (3) a chicken game (which I have also written about before, see here). They also asked the research participants about their beliefs about what the other research participants would choose.

Now, you might be thinking that CEOs have good strategic minds, and so they should be able to do well in game theoretic settings. You might also think that CEOs would be more selfish and more aggressive in these games. In those two hypotheses, you would only be partly correct. Holm et al. find that:
...substantial differences in behavior between the CEOs and the control group, but not in the way many would expect. The CEOs were not in general closer to the Nash equilibrium prediction (assuming selfish preferences). On the contrary, the average control group behavior was closer to the Nash equilibrium in the majority of the games and did not best respond less frequently to their beliefs. The most striking and consistent pattern was that the CEOs had higher expected earnings than the comparison group in all the games. The CEOs cooperated more and played less hawkishly compared to the control group, no matter how the game was framed (abstractly or with a narrative). Compared to the control group the CEOs’ also had significantly higher average beliefs that others would cooperate in the Prisoner’s Dilemma.
More specifically, the CEOs were between 13 and 25 percentage points more likely to choose the cooperative (prisoners' dilemma) or less aggressive (battle of the sexes, or chicken) option that the control group of professionals were. Because of (or perhaps in spite of) this, they earned more overall in the games. So, it appears that CEOs really do act differently than other (otherwise similar) people. Just not in the way that we might expect.

Holm et al. argue that this may be because less aggressive choices may be helpful because they allow the CEO "to mobilize support and loyalty from employees and business partners". I think we would need a lot more research before we can draw any conclusions about the mechanisms that explain these observed differences. Hopefully, there is more research on this to come.

[HT: Marginal Revolution, last year]

Tuesday, 28 May 2019

Autonomous cars, pedestrians, and the game of chicken

I was interested to read this article in The Conversation last month by Jason Thompson (University of Melbourne) and Gemma Read (University of the Sunshine Coast), about the interactions of humans and autonomous cars. In particular, they use some very simple game theory to represent the interaction between cars (autonomous or otherwise) and humans:
A simple example of how this might happen comes from game theory. Take two scenarios at an intersection where pedestrians and vehicles negotiate priority to cross first. Each receives known “pay-offs” for behaviour in the context of the other’s action. The higher the comparative pay-off for either party, the more likely the action.
In the left-hand scenario below, the Nash equilibrium (the optimum combined action of both parties) exists in the lower left quadrant where the pedestrian has a small incentive to “stay” to avoid being injured by the manually driven car, and the driver has a strong incentive to “go”.
However, in the scenario on the right, the autonomous vehicles has a desire to act flawlessly and pose no threat to the pedestrian at all. While this might be great for safety, the pedestrian can now adopt a strategy of “go” at all times, forcing the AV to stay put.
Here's their associated diagram. The red explosions show the Nash equilibrium in each case:

However, I think they have their analysis wrong. I don't think they've really considered the game theory implications here fully. In the game on the left, the car has a dominant strategy to "Go" - going provides a payoff that is always higher than staying. The car driver should never choose to stay, including if the pedestrian chooses to "Go" as well. But, if both pedestrian and driver choose to "Go", then the car runs over the pedestrian. It's hard to see how the car driver is better off going if the pedestrian goes (unless the disutility of washing the blood off their car really is less than the utility gained from getting through the intersection faster).

Similarly, in the game on the right, the pedestrian has a dominant strategy to "Go". But again, would they really choose to "Go" if the car is going too? Maybe in Thompson and Read's world, pedestrians aren't seriously injured or killed when they run into cars, but in the real world it seems like a pedestrian would be pretty stupid to go if a car is going.

A more realistic representation of the game is in the payoff table below. If both the car and the pedestrian go, then both lose. The loss to the car driver is pretty high (damage to their car; maybe jail time for running down a pedestrian), but not as high as the loss to the pedestrian (being injured or killed by a car is pretty serious). If car or pedestrian goes, and the other doesn't, then whichever one goes gets a larger positive payoff. If both car and pedestrian stay, then a standoff ensues, and both get a payoff of zero.


Where is the Nash equilibrium in this game? We can use the 'best response method' to find the equilibrium. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the textbook definition of Nash equilibrium). In this game, the best responses are:
  1. If the car chooses to go, the pedestrian's best response is to stay (since 1 is a better payoff than -5) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If the car chooses to stay, the pedestrian's best response is to go (since 3 is a better payoff than 0);
  3. If the pedestrian chooses to go, the car's best response is to stay (since 1 is a better payoff than -3); and
  4. If the pedestrian chooses to stay, the car's best response is to go (since 3 is a better payoff than 0).
There are two Nash equilibriums in this game: (1) where the car goes and the pedestrian stays; and (2) where the pedestrian goes and the car stays. Notice that both players in this game would be better off if they are the one to go, and worse off if they are the one to stay. But, if both choose to go, then they end up in the worst possible outcome! This is an example of the chicken game (which I have previously blogged about here).

A game with two Nash equilibriums doesn't really give us much guidance as to what the ultimate outcome will be. Both players want to go, but both want to avoid the situation where they are both trying to go at the same time. So, how do we solve this problem?

Currently, we solve the problem using a mixture of road rules and social norms. At pedestrian crossings, cars always stay and pedestrians can go. At traffic lights, cars and pedestrians stay or go depending on whether their light is red or green. At other places, pedestrians stay and cars go - that's why we learn as children to look both ways before crossing the road (a social norm).

However, autonomous cars create an interesting situation, as Thompson and Read highlight in their article. If autonomous cars are programmed to always give way to pedestrians (or, equivalently, to always avoid a collision with a pedestrian where possible), then that means the car will always stay. If the pedestrian knew (for sure) that the car was going to stay, they would choose to go. Every time. Suddenly, we end up in the top right box of the payoff table.

Why is this a problem? As Thompson and Read note:
Now imagine crossing a road or highway in a city saturated by autonomous cars where the threat of being run over disappears. You⁠ ⁠(⁠o⁠r⁠ ⁠a⁠n⁠y⁠ ⁠o⁠t⁠h⁠e⁠r⁠ ⁠mildly i⁠n⁠t⁠e⁠l⁠l⁠i⁠g⁠e⁠n⁠t⁠ ⁠a⁠n⁠i⁠m⁠a⁠l⁠)⁠ might quickly learn that ⁠⁠oncoming traffic poses no threat at all. Replicated thousands of times across a dense inner city, this could produce gridlock among safety-conscious autonomous vehicles, but virtual freedom of movement for humans – maybe even heralding a return to pedestrian rights of yesteryear.
Thompson and Read make the (obviously tongue-in-cheek) suggestion that a solution is for autonomous cars to occasionally, purposefully, run into pedestrians. That seems unlikely (though potentially effective).

Ultimately, this game of chicken is one that pedestrians are likely to win. So, while many are claiming that autonomous vehicles are the solution to traffic problems, they could actually end up making them worse.

Sunday, 23 September 2018

Is Trump vs. Xi a game of chicken, or a prisoners' dilemma?

There are several famous games that we use to teach game theory, one of which is the prisoners' dilemma (which I blogged on earlier in the week). Another is the game of chicken.

In the classic chicken game, two rivals line their cars up at opposite ends of the street. They then race directly towards each other. Each rival then has two options: (1) to swerve out of the way; or (2) to speed on. If one rival swerves and the other speeds on, the rival that sped on wins and the other driver looks foolish and loses some street cred. If both swerve, they both look a bit foolish. If both speed on, then there is a horrific accident and both may be severely injured or die. The game is presented in the payoff table below, for two drivers (Driver A and Driver B).


To find the Nash equilibrium in this game, we use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium). In this game, the best responses are:
  1. If Driver A speeds ahead, Driver B's best response is to swerve (since a loss of face is better than dying in a fiery crash - maybe not immediately, but certainly in the long term) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If Driver A swerves, Driver B's best response is to speed ahead (since winning is better than looking a little foolish);
  3. If Driver B speeds ahead, Driver A's best response is to swerve (since, again, a loss of face is better than dying in a fiery crash);
  4. If Driver B swerves, Driver A's best response is to speed ahead (since winning is better than looking a little foolish).
Notice that there are two Nash equilibriums in this game - where one driver swerves and the other speeds ahead. Both drivers prefer the outcome where they are the one speeding ahead though, so if both try to get that outcome, we end up with both drivers dying in a fiery crash. The chicken game suggests that both players, acting in their own selfish best interest (or not considering the response of the other driver), leads to the worst possible outcome.

Which brings me to this article by Ambrose Evans-Pritchard in the Telegraph UK (gated, but there is an ungated version here):
The US and China are on a combustible escalation path that can end only when there is economic blood on the floor and the political pain threshold of one side or the other has been hit.
Both think they can withstand the longer siege. Neither can retreat easily...
China has in any case stated already that it will match each round of US tariffs with a riposte in kind.
Beijing must carry out this threat or lose face, and Trump has already vowed to escalate further when it does.
It is very hard to see how asset markets priced for perfection can ignore this deranged game of chicken for much longer. The mystery is that they have not crumbled yet.
Is this really a game of chicken? It turns out that it depends on how you define the payoffs. There are two players in the game (the U.S. and China), and two strategies (enact tariffs or hold off - the equivalents of speeding ahead or swerving). So, we can represent the game easily in a payoff table.

First, let's consider the payoffs to each country. If one country enacts tariffs and the other holds off, both countries are worse off than the status quo (they both lose some gains from trade), but the country that holds off probably loses less (their exporters are a bit worse off) [*]. We'll say that the payoff to the country enacting the tariffs is "bad", but for the other country the payoff is just "not so bad". If both countries enact tariffs then all of the bad stuff happens (exporters are a bit worse off, and there are lost gains from trade). We'll say that payoff is "very bad" for both countries. If both countries hold off, then the status quo prevails - the payoff is "OK" for both countries. The game is presented in the payoff table below.


Again solving for Nash equilibrium, the best responses are:
  1. If China enacts tariffs, the U.S.'s best response is to hold off (since "not so bad" is better than "very bad");
  2. If China holds off, the U.S.'s best response is to hold off (since "OK" is better than "bad");
  3. If the U.S. enacts tariffs, China's best response is to hold off (since "not so bad" is better than "very bad");
  4. If the U.S. holds off, China's best response is to hold off (since "OK" is better than "bad").
Notice that the best response for both countries is to hold off, regardless of what the other country does. Holding off is a dominant strategy. There is one Nash equilibrium here, which is for both countries to hold off (the status quo). It is also a dominant strategy equilibrium (because both countries have a dominant strategy). The equilibrium outcome of this game is the best outcome overall.

Clearly, that isn't the game that is playing out at the moment though, so how are things different? The current game is not a game about trade, it is a game about political posturing. The players are not the U.S. and China, but Donald Trump and Xi Jinping. They want to look strong, and not appear weak (to each other, or to their respective peoples). So, the payoffs and the players are different. The game that is actually being played looks more like the payoff table below. If both hold off, the status quo prevails (the payoff is "OK" for both. If one of them enacts tariffs and the other holds off, whichever of them enacted tariffs appears strong, and the other appears weak. If both enact tariffs, the payoff is costly to the economy.


Again solving for Nash equilibrium, the best responses are:
  1. If Xi enacts tariffs, Trump's best response is to enact tariffs (since "costly" is better than "weak");
  2. If Xi holds off, Trump's best response is to enact tariffs (since "strong" is better than "OK");
  3. If Trump enacts tariffs, Xi's best response is to enact tariffs (since "costly" is better than "weak");
  4. If Trump holds off, Xi's best response is to enact tariffs (since "strong" is better than "OK").
Notice that the best response for both leaders is to enact tariffs, regardless of what the other leader does! Enacting tariffs is a dominant strategy, and both leaders enacting tariffs is both the only Nash equilibrium and a dominant strategy equilibrium. The single equilibrium is also unambiguously worse than one of the other outcomes - this is an example of the prisoners' dilemma. Notice that it is not a chicken game.

How could this become a chicken game? If costing the economy was worse than appearing weak, then that would change things around. In that case, the best response to the other leader enacting tariffs would be to hold off. There would be two Nash equilibriums - where one leader enacts tariffs and the other holds off. However, both would prefer to be the leader enacting the tariffs rather than the one holding off.

So, whether this game is a chicken game or a prisoners' dilemma depends on how you think each leader feels about appearing weak. It seems to me that both want to avoid that at all costs. In my mind, this is a prisoners' dilemma, not a chicken game.

The repeated prisoners' dilemma can be solved for the optimal outcome (both holding off), but this requires cooperation between the two leaders. In order for this cooperation to arise, each leader must trust the other (because enacting tariffs is still a dominant strategy). If we want global trade to survive this showdown, somehow we need these leaders to develop a trusting relationship. It's a pity that Trump will not be at the APEC leaders meeting - it seems like a group hug is in order!

*****

[*] This might sound surprising. For the country imposing tariffs, tariffs lead to a deadweight loss (lost wellbeing). They make domestic sellers better off, but make domestic consumers worse off by more than the gain to domestic sellers. In contrast, the market in the country that holds off has no deadweight loss. The exporting firms in that country will be able to export a bit less, but that probably doesn't have as big of a negative impact as the tariffs do on the country that imposed them.

Wednesday, 15 February 2017

Brexit negotiations as a game of chicken

In a post last month, Tim Harford perceptively characterised the posturing between Britain and the European Union over Brexit as a game of chicken:
First: to be an effective negotiator often means accepting some risk of disaster. The simplest model of this is the game of “Chicken”, in which two leather-clad rebels get into their cars, and drive towards each other at a furious pace. The first one to veer off the road loses his dignity, unless neither of them swerve, in which case both of them will lose a lot more than that.
Chicken is an idiotic game, whose players have little to gain and much to lose. But Chicken teaches us that you can gain an advantage by limiting your own options. Imagine detaching your steering wheel and flamboyantly discarding it as you race headlong towards your opponent. Victory would be guaranteed. Nobody would drive straight at a car that cannot steer out of the way. But here’s a worrisome prospect: what if, as you hurl your own steering wheel out of the window, you notice that your rival has done exactly the same thing?
All this matters because both the UK and the EU are doing their best to give the impression that they’ve thrown their steering wheels away. Control of immigration is non-negotiable, says Theresa May. Fine, says the EU — in that case membership of the single market is out of the question. Fine, says May: we’re out. Don’t let the door hit you as you leave, says the EU.
It’s easy to see why both sides are behaving like this — it’s the logic of Chicken. But the eventual result may be something no sane person wants: a car crash. In May’s recent speech, she set out her willingness to risk such a crash by saying she might walk away without a deal. That does make some sense: it’s how you act if you want to win a game of Chicken. But there are games of Chicken that nobody wins.
The Brexit negotiations chicken game is laid out in the table below. The EU and the UK can choose to 'make concessions', or to 'play hardball'. If both make concessions, the outcome is essentially pretty neutral (a payoff of zero for both of them). However, if either the EU or the UK plays hardball while the other makes concessions, whichever of them plays hardball comes out better off (positive payoff) at the expense of the other (negative payoff).  Finally, if both play hardball, both will be much worse off (very negative payoffs).


Where are the Nash equilibriums in this game? To identify them, we can use the 'best response' method. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium).

For our game outlined above:
  1. If the UK makes concessions, the EU's best response is to play hardball (since + is better than 0) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If the UK plays hardball, the EU's best response is to make concessions (since - is better than --);
  3. If the EU makes concessions, the UK's best response is to play hardball (since + is better than 0); and
  4. If the EU plays hardball, the UK's best response is to make concessions (since - is better than --).
Note that there are two Nash equilibriums, where one of the EU or the UK plays hardball, and the other makes concessions. However, both of them want to be the one playing hardball. This is a type of coordination game, and it is likely that both the EU and the UK will try to play hardball (but leading both to incur big losses!).

The solution to getting your preferred equilibrium outcome in the chicken game is to make a credible commitment (such as removing the steering wheel that Harford suggests for the classic chicken game). In the case of Brexit though, it isn't clear how either side can make a credible commitment to the hardball strategy, and both are already moving their feet towards the accelerator. But if neither are willing to make concessions, the outcome is clear.

Read more:


Saturday, 19 December 2015

China's zombie companies are playing chicken

Earlier this month, Andrew Batson wrote an interesting blog article about China's zombie companies:
One of the more interesting developments in official Chinese discussions about the economy has been the appearance of the term “zombie companies”... money-losing companies that seem to stay alive far longer than economic fundamentals warrant. This problem is particularly acute in the commodity sectors: a global supply glut has driven down prices of iron ore and coal to multi-year lows, levels where China’s relatively low-quality and high-cost mines have difficulty being competitive. And yet they continue operating despite losing money, because it is easier to keep producing than to completely shut down.
In ECON100 we no longer cover cost curves in detail, so we also don't talk about the section of the firm's marginal cost curve where it makes losses but prefers to continue trading because the losses from trading are smaller than the losses from shutting down. However, this is exactly the situation for China's zombie firms. As Batson notes:
An excellent story this week in the China Economic Times on the woes of the coal heartland of Shanxi quoted one executive saying, “If we produce a ton of coal, we lose a hundred yuan. If we don’t produce, we lose even more.”
Another aspect of the reluctance of China's zombie companies to shut down is strategic, and in this case it may be that the zombie companies continue to operate even if their losses would be smaller by shutting down. What the zombie companies are doing is playing a form of the 'chicken game'. In the classic version of the game of chicken, the two players are driving cars and line up at each end of the street. They accelerate towards each other, and if one of the drivers swerves out of the way, the other wins. If they both swerve, neither wins, and if neither of them swerve then both die horribly in a fiery car accident.

Now consider the game for zombie companies, as expressed in the payoff table below (assuming for simplicity that there are only two zombie firms, A and B). If either firm shuts down, they incur a small loss (including if both firms shut down). However, if either firm continues operating while the other firm shuts down, the remaining firm is able to survive and return to profitability. Finally, if both firms continue operating, both incur a big loss.


Where are the Nash equilibriums in this game? To identify them, we can use the 'best response' method. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium).

For our game outlined above:
  1. If Zombie Company A continues operating, Zombie Company B's best response is to shut down (since a small loss is better than a big loss) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If Zombie Company A shuts down, Zombie Company B's best response is to continue operating (since survival and profits is better than a small loss);
  3. If Zombie Company B continues operating, Zombie Company A's best response is to shut down (since a small loss is better than a big loss); and
  4. If Zombie Company B shuts down, Zombie Company A's best response is to continue operating (since survival and profits is better than a small loss).
Note that there are two Nash equilibriums, where one of the companies shuts down, and the other continues operating. However, both firms want to be the firm that continues operating. This is a type of coordination game, and it is likely that both firms will try to continue operating, in the hopes of being the only one left (but leading both to incur big losses in the meantime!).

What's the solution? To avoid the social costs of the zombie companies continuing to operate and generating large losses, the government probably needs to intervene. Or, as noted at the bottom of the Batson blog article, mergers of these firms will remove (or mitigate) the strategic element, which is the real problem in this case.