Showing posts with label Game theory. Show all posts
Showing posts with label Game theory. Show all posts

Monday, 11 May 2026

David Oks on the bad business economics of airlines

Airlines are a strange business. They seem to have huge numbers of passengers, and yet we routinely hear about airlines struggling financially, entering administration, or shutting down. For example, in 2024 in Australia, Bonza collapsed, while Rex withdrew from major intercity routes after entering vountary administration (see this post for more on those examples). The most recent example is Spirit Airlines in the US, which shut down this month, after its second bankruptcy process in less than two years.

I just discovered David Oks's Substack, which has overnight become one of my favourites. The post that first attracted me was this one titled 'Why airlines are always going bankrupt', inspired by the Spirit Airlines story. Another recent post titled 'Why ATMs didn’t kill bank teller jobs, but the iPhone did' is also excellent, as is 'How funerals keep Africa poor'.

Oks's airline post has a huge amount of depth, so is difficult to summarise without losing part of the story. However, I'm going to try (but I really recommend that you read his entire post, as it is truly excellent).

Oks first notes that airlines are not just badly managed, they are structurally vulnerable and often fail to earn a positive return on capital. He then turns to the game theory of airlines, noting that there is an 'empty core' problem, meaning that there isn't any subset of the airline industry that can form a stable coalition, because some part of the coalition would always be able to make themselves better off by breaking away from the coalition. In other words, no group of airlines and routes is stable for long, because whenever capacity is tight and fares are high, another airline has an incentive to add seats. However, once those seats are added, the market can quickly become unprofitable.

The 'empty core' arises because airlines have high fixed costs, low marginal costs, volatile demand, weak product differentiation, and large minimum efficient scale. All of that means that adding one extra airline on a route, or one extra aircraft, can swing a market from being undersupplied and profitable to being undersupplied and unprofitable. So, what happens in the airline industry is that when there are few airlines, they are profitable and happy, but that encourages new airlines to enter, after which profits decrease until one or more of the airlines shuts down. And then the cycle starts over.

As you can see, there is a lot to unpack from Oks's post. Again, I encourage you to read the whole post. But the key points are, first, that airlines are always going bankrupt because the nature of the airline industry lends itself to a boom-bust cycle when airlines compete with each other.

Second, from the perspective of airline consumers (and, probably, governments as well) there is an uncomfortable trade-off. We want low fares from competition between airlines, and we also want airlines to be financially stable, but it may be difficult to have both. The pursuit of low fares can set off the cycle described above, with competition pushing fares down, leading to falling profitability, and eventually one or more airlines exiting the market. Airline financial stability, on the other hand, probably requires some limit on competition between airlines. One way of achieving this is to regulate airlines more closely, closer to the equilibrium that existed before deregulation in the late 1970s and early 1980s, when regulators had much greater control over fares, routes, and market entry. Airlines segmented the market, flew fewer routes, and were profitable in part because they were not actively competing with each other. The other alternative is to recognise that airlines can diversify into other markets. Oks’s most striking example of this is Delta:

...the most profitable airline in the United States, which started a fruitful partnership with American Express in 1996 and launched a co-branded card with them in 2008. Annual spending on Delta-branded American Express cards comes out to about 1 percent of U.S. GDP. In 2025, this produced about $8 billion in revenue for Delta, accounting for more than the entirety of its profit. That means that without the American Express partnership, Delta would be operating at a substantial loss. In effect, Delta’s aviation business is a loss leader for a much more profitable credit card partnership. So to the extent that Delta is now a good business, it is because it escaped the basic instability of the airline industry by becoming less of an airline.

Allowing airlines to operate in a cartel-like equilibrium, as airlines effectively did prior to the 1980s, doesn't strike me as a very positive solution, especially from the perspective of consumers. Having airlines shut down at short notice, cancelling thousands of flights, is clearly not good for consumers either. Perhaps then we need to tolerate airlines that try to sell us financial services, or unbundling the airline package and separately selling seat selection, checked baggage, meals, and so on (see this post about Jetstar's business practices, or this one).

Saturday, 9 May 2026

The economics of castles

When I'm in Britain or Ireland, one of my favourite sightseeing trips is to visit medieval castles. Even the ruined ones are fun to visit. Actually, maybe the ruined ones are more fun to visit, because you get to imagine what they would have looked like in their heyday. Britain and Ireland are full of castles, many of which were built by and housed local nobles. In fact, in relative terms there were very few royal castles, which the literature in history and economic history has interpreted as a sign that centralised states were weak.

However, this recent article by Desiree Desierto and Mark Koyama (both George Mason University), published in the journal European Economic Review (ungated earlier version here) challenges that view. They instead show that there was an economic logic to the proliferation of private castles.

Desierto and Koyama develop a game theoretic model of medieval states, which first recognises that the monarch cannot rule alone but must rely on a coalition of local lords or barons. Each lord agrees to join the coalition, and pledges resources to the monarch in exchange for a (however small) share of control of the kingdom. The monarch can renege on the agreement, taking the resources without offering a share to the lord. However, the lord would then rebel, leaving the coalition. What allows the lord to leave the coalition, and gives them bargaining power, is the presence of their own castles, since they can retreat to their castle when they rebel against the monarch. Without a private castle, the lord would have little bargaining power, and would anticipate the monarch reneging on any agreement, and so they would not join the coalition in the first place. In economic terms, the castle gives the lord an outside option

So, the monarch tolerates private castles held by the lords, because the presence of those castles gives the lords the feeling of security they need to join the kingdom. And, in turn, the presence of those castles disciplines the monarch. Rebellions are more costly to suppress when the lords can retreat to a well-defended private castle. The lords' outside option increases the feasibility of rebellion and ensures that the monarch mostly keeps to their agreement with the lords.

In short, the private castles induce an equilibrium where the kingdom is larger and more stable than it would be without them. Notice that this is the opposite of the conventional view that private castles represent a sign that a state was weak.

Desierto and Koyama support their argument with descriptive evidence, noting that:

In Norman England after the Conquest, castles were built across the country: by 1154 there were 225 baronial castles (compared to 49 royal castles) in England... Baronial castles allowed the Dukes of Normandy to extend their authority over the far larger territory of Anglo-Saxon England. Similarly, in the twelfth century Angevin rule expanded over much of France as semi-independent lords in Gascony accepted the lordship of Henry II.

Desierto and Koyama also note that powerful medieval monarchs did not act to systematically eliminate private castles, and in fact the monarchs often gave away their own castles to local lords. And the power of the lords did keep the monarchs in check - a lord's probability of rebelling against the monarch was positively correlated with the number of private castles in the lord's family network. That last point might sound contradictory, but since castles make rebellion by lords a more credible threat, this can deter monarchs from reneging and reduce the number of rebellions overall. However, when rebellion does occur, lords connected to more castles were more likely to participate in the rebellion.

So, if private castles were so important for the stability of medieval states, why did private castles eventually disappear? Desierto and Koyama note that:

The answer is military technology, not the rising power of the state. Technological changes and the associated ‘‘military revolution’’ that took place beginning in the late Middle Ages reduced the value of medieval fortifications. The main technological innovation was the introduction and improvement of gunpowder weapons, which began in the fourteenth century but only really began to have a serious impact in the fifteenth century with the introduction of iron cannonballs.

This is not a new insight, but it does align well with their model. However, it somewhat reverses the logic of the conventional view, which is that greater state power, along with military technology, reduced the prevalence of private castles. Instead, in Desierto and Koyama's model, the rise of gunpowder reduces the ability of lords to retreat to a well-defended castle in the event of a rebellion (because the castle could not be as well-defended against cannons). This reduced the lords' bargaining power, giving the monarch and the centralised state greater power. As a result, the state should become less stable. In support of this, Desierto and Koyama use the Wars of the Roses in England as an example:

England experienced a large number of rebellions and civil wars between 1450 and 1500. These conflicts are conventionally grouped under the label of the Wars of the Roses (1455–1485), but the period of weak state capacity and frequent rebellion extended from Jack Cade’s uprising in 1450 through Perkin Warbeck’s invasion and the Second Cornish Uprising in 1497. The causes of these rebellions were complex, multifaceted, and varied across cases. Nonetheless, the frequency of civil war during this period is consistent with our model’s prediction that a decline in the military value of castles would destabilize feudal realms.

And so, as gunpowder reduced the military value of castles, private castles became much less useful as a source of bargaining power for lords. That helps explain why the medieval pattern of widespread private castles gave way to state-controlled castles from the mid-15th Century onwards. Now, I'll be thinking more carefully about the vintage of the castles I visit on my next trip to Europe next month!

Tuesday, 24 March 2026

Evidence that artificial intelligence is increasing the impact, but narrowing the scope, of research

There is growing evidence of positive impacts of generative artificial intelligence on productivity. This includes productivity in research (see this post, for example), including my own. However, some have questioned whether increasing research productivity comes at a cost of narrowing the scope of research.

So, I was interested to read this article by Qianyue Hao (Tsinghua University) and co-authors, published in the prestigious journal Nature (ungated earlier version here) late last year. They look at the impact of AI tools (not limited to generative AI) on the productivity of researchers and the quality of research. Specifically, they look at authors publishing in six representative fields: biology, medicine, chemistry, physics, materials science, and geology, across three 'eras': (1) the 'machine learning era ' (from 1980 to 2014), the 'deep learning era' (from 2015 to 2022), and the 'generative AI era' (from 2023 onwards). Hao et al. compare authors who publish 'AI augmented papers' with those who do not. An 'AI augmented paper' is one that uses methods such as:

...support vector machines and principal component analysis from the machine learning era, and convolutional neural networks and generative adversarial networks from the deep learning era. Large language models, which have emerged in recent years, also rank among the most frequently used methods...

Using a dataset that includes over 27 million papers with complete records that were published between 1980 and 2025, of which about 310,000 were 'AI augmented', Hao et al. find that:

...annual citations to AI papers are 98.70% higher than those to non-AI papers on average...

So, AI augmented research gathers more citations, which suggests that authors using AI in their research achieve greater impact. This is reinforced by evidence that AI augmented papers are published in higher quality journals (with Q1 journals being the highest ranked). Hao et al. report that:

...the proportion of AI papers in Q1 journals is 18.60% higher than that of non-AI papers in all journals; in Q2 journals, the AI proportion is 1.59% higher; whereas Q3 and Q4 journals hold a relatively lower proportion of papers with AI... These results indicate a heterogeneous distribution of AI-augmented papers across journals, with a higher prevalence in high-impact journals.

And AI appears to make authors more productive, as:

On average, researchers adopting AI annually publish 3.02 times more papers... and garner 4.84 times more citations... than those not adopting AI, with consistency.

All of these results seem to hold across all of the disciplines that Hao et al. consider. However, it is not all good news. Hao et al. use machine learning to create a measure of the 'breadth of scholarly attention'. Using that measure, they find that:

Compared with conventional research, AI research is associated with a 4.63% contracted median collective knowledge extent across science, which is consistent across all six disciplines... Moreover, when dividing these disciplines into more than two hundred sub-fields, the contraction of knowledge extent can be observed in more than 70% of them...

Of course, some of the differences here may be due to selection, as the types of researchers, and the types of research, involving AI use may be meaningfully different from those that don't. However, putting the selection issues aside, Hao et al. note that there is a tension between the individual researcher's incentive to produce a greater quantity of research that has higher impact, which would suggest greater use of AI, and the social incentive to produce a greater breadth of research.

So, the takeaway from this paper is that we need to consider researcher incentives, not just productivity. Specifically, this research suggests that the use of AI in research is leading to a 'prisoners' dilemma' outcome: each individual researcher acting in their own best interests (and using AI in their research) leads to an outcome that is worse for society overall (less breadth of research and more incremental gains).

Hao et al. conclude that:

The substantial academic benefits of AI use may be a driving force behind its accelerated rate of adoption; however, we also find unintended consequences from the increased prevalence of AI-augmented research. In all fields, AI-augmented research focuses on a narrower scope of scientific topics and reduces the scientific engagement of follow-on research, leading to more overlapping research work that slows the expansion of knowledge. Further, with a greater concentration of collective attention to the same AI papers, the adoption of AI seems to induce authors to converge on the same solutions to known problems rather than create new ones.

So, what is the solution here? Society probably wants research to be higher quality and have a broad scope. But individual researchers' incentives to use AI in their research appears inconsistent with that outcome. The traditional prisoners' dilemma is a repeated game (see here or here, for example), and the players of that game can avoid the worst outcome by cooperating. In this case, the researchers could cooperate by agreeing not to use AI in their research. The problem is that every researcher has an incentive to cheat on that agreement, since if they use AI, then that will be good for their career. This prisoners' dilemma is more difficult to ensure cooperation in than the traditional game, because there are not just two players who need to cooperate, but thousands (or millions). Ensuring cooperation in a prisoners' dilemma game with many players, each of whom is far better off cheating than cooperating, is almost impossible (which is why solving the problem of climate change is so difficult).

My own view is that the answer is not to keep AI out of research. That is not realistic, in the same way that it's not realistic to expect students not to use generative AI. The incentives need to be redesigned, but this will be no easy task. As long as universities, research funders, and publishers reward researchers for quantity, citations, and publication in top-ranked outlets, then we should expect more AI-augmented work, with a narrower scope than society might prefer. If we want AI to expand knowledge rather than simply accelerate competition within narrow foci, then we need institutions that also reward novelty, breadth, and the discovery of new questions. That is the economic challenge we must face up to.

[HT: Marginal Revolution]

Tuesday, 2 September 2025

The economics of pricing LLM tokens

Ethan Ding had a really interesting post on Substack last month, discussing his view on the future costs of tokens for large language models (LLMs), and what that means for the viability of subscription-based generative AI. I want to focus on two aspects of Ding's post. First, this (the lack of capitalisation is Ding's style):

the math has fundamentally broken.

prisoner’s dilemma for everyone else

this leaves everyone else in an impossible position.

every ai company knows usage-based pricing would save them. they also know it would kill them. while you're being responsible with $0.01/1k tokens, your vc-funded competitor offers unlimited for $20/month.

guess where users go?

classic prisoner's dilemma:

  • everyone charges usage-based → sustainable industry
  • everyone charges flat-rate → race to the bottom
  • you charge usage, others charge flat → you die alone
  • you charge flat, others charge usage → you win (then die later)

so everyone defects. everyone subsidizes power users.

Ok, so let's look at this prisoners' dilemma. Consider two AI firms (Firm A and Firm B), each with two strategies to choose from (Usage-based pricing, or flat-rate pricing). The game is outlined in the payoff table below. The payoffs are expressed in +'s and -'s, with more +'s obviously being better.

To find the Nash equilibrium in this game, we use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium). In this game, the best responses are:

  1. If Firm B chooses usage-based pricing, Firm A's best response is to choose flat-rate pricing (since ++ is a better payoff than +) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If Firm B chooses flat-rate pricing, Firm A's best response is to choose flat-rate pricing (since - is a better payoff than --);
  3. If Firm A chooses usage-based pricing, Firm B's best response is to choose flat-rate pricing (since ++ is a better payoff than +); and
  4. If Firm A chooses flat-rate pricing, Firm B's best response is to choose flat-rate pricing (since - is a better payoff than --).

Note that Firm A's best response is always to choose flat-rate pricing. This is their dominant strategy. Likewise, Firm B's best response is always to choose flat-rate pricing, which makes it their dominant strategy as well. The single Nash equilibrium occurs where both players are playing a best response (where there are two ticks), which is where both firms choose flat-rate pricing.

Notice that both players would be unambiguously better off if they chose usage-based pricing. However, both will choose flat-rate pricing, which makes them both worse off. This is a prisoners' dilemma game (it's a dilemma because, when both players act in their own best interests, both are made worse off).

Ding notes that this is a losing proposition for all generative AI firms, and is the position that they are all in right now. They could try to cooperate and shift to usage-based pricing, but there will always be a strong incentive for the firms to cheat on any agreement and instead offer flat-rate pricing. So, any agreement will not last. Especially since there are other strategies available, which Ding goes on to discuss. The one that caught my eye was this:

use ai as a loss leader to drive consumption of aws-competitive services. you're not selling inference. you're selling everything else, and inference is just marketing spend.

the genius is that code generation naturally creates demand for hosting. every app needs somewhere to run. every database needs management. every deployment needs monitoring. let openai and anthropic race inference to zero while you own everything else.

the companies still playing flat-rate-grow-at-all-costs? dead companies walking. they just have very expensive funerals scheduled for q4.

It makes sense to play the losing prisoners' dilemma strategy, if a firm can use it to be more profitable elsewhere. Using generative AI as a loss leader, and then making more profits by selling complementary services (hosting, data management, monitoring) may be more profitable overall for the generative AI firms.

For loss leading to be successful though, two conditions need to be met. First, the loss leading service should be price elastic. That means that when price is low, many consumers are attracted to the service. That seems likely to be the case for generative AI, because when the price increases, consumers can easily switch to one of the many other generative AI platforms. Second, there must be many other complementary services for the firm to sell. The three suggestions by Ding (hosting, data management, monitoring) are all complements to generative AI (or, at least, to the ways that generative AI is being used right now). So, it seems that loss leading with generative AI may be a profitable strategy for the generative AI firms, even though it means playing out the prisoners' dilemma on pricing.

[HT: Marginal Revolution]

Saturday, 14 June 2025

Penalty shootouts and first-mover advantage

I enjoyed watching the UEFA Nations League final on Monday. Spain and Portugal put on a good show, and the scores were tied at 2-2 at the end of extra time. The game went to a penalty shootout. Portugal had the first penalty shot, and eventually ended up winning the shootout 5-3, after Portuguese goalkeeper Diogo Costa saved a weak shot by Alvaro Morata.

Would the result have been different if Spain had taken the first penalty shot? There certainly is conventional wisdom that says that going first in a penalty shootout conveys an advantage (a first-mover advantage in game theory terminology). The argument is that, because the team going second is often trying to come from behind, that team faces more pressure than the team going first.

However, the evidence in favour of that conventional wisdom has been challenged, most recently and most thoroughly in this new article by David Pipke (Kiel Institute for the World Economy), published in the Journal of Economic Psychology (open access). Pipke looks at the outcomes of 7116 penalty shootouts from 1970 to 2024, across top leagues and international competitions. He then tests whether the outcome deviates from a random outcome (in which the team kicking first wins 50 percent of the time). He finds that:

In soccer, the first-kicking team wins 50.2 % of the time (p =0.785) across 7,116 matches in the Flashscore data.

So, there is no statistical evidence for a first-mover advantage in penalty shootouts in football (soccer). Pipke then turns to ice hockey, which also features shootouts but where the probability of a successful shot in a shootout is much lower. Using data from 4407 shootouts in North American ice hockey leagues over the period from 2010 to 2024, Pipke finds that:

In ice hockey, the first team wins 48.9 % of shootouts (p =0.148)...

It's closer to statistical significance, but not quite. There is no evidence for a first-mover advantage in ice hockey shootouts either. Pipke then notes that his statistical tests can:

...reject the hypothesis that the first-mover’s winning probability deviates by more than 1.6 percentage points in soccer... and 2.9 percentage points in hockey from a 50:50 split, at a 1 % significance level.

Pipke then looks at some subsets of the football data, and finds that:

In 342 women’s soccer competitions, the first-moving team wins 172 times (50.3 %, p = 0.957). In youth soccer shootouts, the first-kicking team prevails in 130 out of 277 cases (46.9 %, p = 0.336).

Second, between 2017 and 2019, an alternative format, where teams alternate in an A,B,B,A pattern, was tested in various competitions to address concerns about an inherent advantage of kicking first. In 44 shootouts following this sequence, the first-kicking team won 56.8 % of the time (25 shootouts), with no statistically significant deviation from a 50:50 split (p = 0.451).

So, overall, there is no evidence of a first-mover advantage in a penalty shootout (in football or ice hockey). The result may have been different if Spain had gone first in the UEFA Nations League final penalty shootout, but going first wouldn't have given Spain a statistical advantage.

Read more:

Thursday, 3 April 2025

Mobile phone providers and the repeated switching costs game

This week, my ECONS101 class covered pricing and business strategy, and one aspect of that is switching costs and customer lock-in. Switching costs are the costs of switching from one good or service to another (or from one provider to another). Customer lock-in occurs when customers find it difficult (costly) to change once they have started purchasing a particular good or service. The main cause of customer lock-in is, unsurprisingly, high switching costs.

As one example, consider this article from the New Zealand Herald last month:

A new Commerce Commission study has found the switching process between telecommunications providers is not working as well as it should for consumers...

The study found 50% of mobile switchers and 45% of broadband switchers ran into at least one issue when switching.

The experience was so bad that 29% of mobile switchers and 27% of broadband switchers said they wouldn’t want to switch again in future...

The commission’s latest consumer satisfaction report found that 31% of mobile consumers and 29% of broadband consumers have not switched because it requires ‘too much effort to change providers’...

Gilbertson said a lack of comprehensive protocols between the “gaining” service provider and the “losing” service provider was a central issue with the current switching process.

This led to a number of problems, including double billing, unexpected charges, and delays.

The difficulty of changing from one mobile phone provider to another is a form of switching cost. It's not a monetary cost, but the time, effort, and frustration experienced by consumers wanting to switch makes the process of switching costly. And because the process is costly, mobile phone consumers are locked into their current provider.

It is clear why a mobile phone provider would want to make it difficult (costly) for its consumers to switch away from it and use some other provider. However, why don't mobile phone providers try to make it easier to switch to using their service instead? Maybe they could have staff whose role is to help consumers to navigate the process of switching to their service. That would allow the mobile phone provider to attract consumers and capture a greater market share. The answer is provided by considering a little bit of game theory.

Consider the game below, with two mobile phone providers (A and B), each with two strategies ('Easy' to switch to, and 'Hard' to switch to). The payoffs are made-up numbers that might represent profits to the two providers.

To find the Nash equilibrium in this game, we use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium). In this game, the best responses are:

  1. If Provider B chooses to make switching easy, Provider A's best response is to make switching easy (since 3 is a better payoff than 2) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If Provider B chooses to make switching hard, Provider A's best response is to make switching easy (since 8 is a better payoff than 6);
  3. If Provider A chooses to make switching easy, Provider B's best response is to make switching easy (since 3 is a better payoff than 2); and
  4. If Provider A chooses to make switching hard, Provider B's best response is to make switching easy (since 8 is a better payoff than 6).

Note that Provider A's best response is always to choose to make switching easy. This is their dominant strategy. Likewise, Provider B's best response is always to make switching easy, which makes it their dominant strategy as well. The single Nash equilibrium occurs where both players are playing a best response (where there are two ticks), which is where both providers make switching easy.

So, that seems to suggest that the mobile phone providers should be making switching to them easier. However, notice that both providers would be unambiguously better off if they chose to make switching hard (they would both receive a payoff of 6, instead of both receiving a payoff of 3). By both choosing to make switching easy, it makes both providers worse off. This is a prisoners' dilemma game (it's a dilemma because, when both players act in their own best interests, both are made worse off).

That's not the end of this story though, because the simple example above assumes that this is a non-repeated game. A non-repeated game is played once only, after which the two players go their separate ways, never to interact again. Most games in the real world are not like that - they are repeated games. In a repeated game, the outcome may differ from the equilibrium of the non-repeated game, because the players can learn to work together to obtain the best outcome.

So, given that this is a repeated game (because the providers are constantly deciding whether to make switching easier or not), both providers will realise that they are better off making switching harder, and receiving a higher payoff as a result. And unsurprisingly, that is what happens, and it doesn't require an explicit agreement between the players - the agreement is 'tacit' (it is understood by the providers without needing to be explicit). Each provider just needs to trust that the other providers will make switching hard (because there is an incentive for each provider to 'cheat' on this outcome). Any instance of cheating (by making switching easier) would be immediately known by the other providers, and the agreement would break down, making them all worse off. So, there is an incentive for all providers to keep switching hard for the consumers. Even a new entrant firm into the market, which might initially make it easy for consumers to switch to them in order to capture market share, would soon realise that they are then better off making switching more difficult (it is not so long ago (2009) that 2degrees was a new entrant in this market).

The Commerce Commission is correct that the difficulty of switching mobile phone providers (the switching cost) keeps consumers with their current provider (customer lock-in). The result is that the mobile phone providers can profit from increasing prices for their lock-in consumers. The only solution to this situation would be to find some way to force a breakdown of the tacit arrangement. Then the market would settle at the equilibrium of all providers making it easy to switch to them. This may be an instance where some regulation is necessary.

Sunday, 16 March 2025

Why tit-for-tat tariffs may not work against Trump

Last week, my ECONS101 class covered game theory. At the end of the final lecture, after we had been covering repeated games and tit-for-tat strategies, a really perceptive student asked me about Trump's tariffs. A lot of the rhetoric about tariffs has been posed in terms of tit-for-tat (see here and here, for example). The student's question got me thinking though, about why a tit-for-tat strategy may not work in this case.

Before we get that far though, we need to think about the tariff game, as outlined in the payoff table below. There are two players: USA and 'Other Country'. Each player has two strategies: high tariffs, or low tariffs (which includes no tariffs). The payoffs are expressed as "+" for good outcomes (and "++" is particularly good), and "--" for bad outcomes (and "--" is particularly bad), while zero is a neutral payoff. 

To find the Nash equilibrium in this game, we use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium). In this game, the best responses are:

  1. If the other country chooses high tariffs, USA's best response is to choose high tariffs (since "0" is better than "--" as a payoff) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If the other country chooses low tariffs, USA's best response is to choose high tariffs (since "++" is better than "+" as a payoff);
  3. If USA chooses high tariffs, the other country's best response is to choose high tariffs (since "0" is better than "--" as a payoff); and
  4. If USA chooses low tariffs, the other country's best response is to choose high tariffs (since "++" is better than "+" as a payoff).

Notice that USA chooses high tariffs no matter what the other country does, high tariffs is a dominant strategy for the USA. Similarly, since the other country chooses high tariffs no matter what the USA does, high tariffs is a dominant strategy for the other country. Both countries will choose to play their dominant strategy (because it is always better than the other strategy, not matter what the other country chooses to do). The outcome where both countries choose high tariffs is the Nash equilibrium in this game (it is also a dominant strategy equilibrium, because both countries have a dominant strategy).

However, I'm sure that you can clearly see that both countries would be better off with low tariffs (since the payoff for each country would be "+", instead of "0"). This game is an example of the prisoner's dilemma (it's a dilemma because, when both countries act in their own best interests, both are made worse off).

However, it is important to remember that this game is a repeated game. It is played more than once, with the same players, and the same strategy choices. When a game is repeated, then the outcome may differ from the equilibrium of the non-repeated game, because the players can learn to work together to obtain the best outcome.

In a repeated prisoners' dilemma game like this, each player can encourage the other to cooperate by using the tit-for-tat strategy. That strategy, identified by Robert Axelrod in the 1980s, works by initially cooperating (low tariffs), and then in each play of game after the first, the player does whatever the other player did last time. So, if the USA chooses high tariffs, then the other country should punish them and choose high tariffs in the next play of the game. And if the USA chooses low tariffs, then the other country should reward them and choose low tariffs in the next play of the game. The tit-for-tat strategy works because it encourages the other player to cooperate. And that is what many people have been expecting of Trump. If you punish the USA for high tariffs by setting your own high tariffs, eventually they will realise their error and start cooperating with low tariffs again.

But there is a problem. The whole edifice of the tit-for-tat strategy assumes that Trump knows that he is playing the game outlined above, where there is an unambiguously better outcome for both countries, if they both choose low tariffs. By his past statements, it is absolutely clear that Trump thinks that trade is a zero-sum game (for example, see here or here).

So, what does the game look like if you believe that trade is zero-sum? Instead of there being gains from the low-tariff/low-tariff outcome, the payoffs become zero, as shown in the payoff table below.

Solving this game for the Nash equilibrium, the best responses are:

  1. If the other country chooses high tariffs, USA's best response is to choose high tariffs (since "0" is better than "--" as a payoff) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If the other country chooses low tariffs, USA's best response is to choose high tariffs (since "++" is better than "0" as a payoff);
  3. If USA chooses high tariffs, the other country's best response is to choose high tariffs (since "0" is better than "--" as a payoff); and
  4. If USA chooses low tariffs, the other country's best response is to choose high tariffs (since "++" is better than "0" as a payoff).

Notice that the game itself doesn't change. Imposing high tariffs is still a dominant strategy for both countries, and the outcome where both countries choose high tariffs is the only Nash equilibrium (and is also a dominant strategy equilibrium).

However, there is an important difference in this game when it is played as a repeated game. There is no incentive for players to cooperate. That's because cooperating results in a payoff of "0", just the same as not cooperating. And because of this, the tit-for-tat strategy would be pointless.

Which seems to be what other countries are finding, when dealing with Trump. Other countries may think they are playing the first game, where a tit-for-tat strategy may get Trump to reconsider. But if Trump thinks they are playing the second (zero-sum) game (and it seems that he does), then the tit-for-tat strategy is simply not going to work.

[HT: Sarah from my ECONS101 class]

Thursday, 13 March 2025

Hawks, doves, Israel and Iran

In The Conversation last October, Andrew Thomas (Deakin University) discussed the recent (at that time) military flare-up between Iran and Israel, likening it to a 'game of chicken':

Israel’s strike on military targets in Iran over the weekend is becoming a more routine occurrence in the decades-long rivalry between the two states...

There is a reason why direct military strikes between nations are rare, even between sworn enemies. When attacking another state, it is difficult to know exactly how they will respond, though a retaliatory strike is almost often expected.

This is because defence forces are not just used for fighting and winning wars – they are also vital to deterring them. When a fighting force is attacked, it’s important for it to strike back to maintain the perception it can deter future attacks and make a display of its capabilities. This is what is happening right now between Israel and Iran – neither side wants to appear weak.

If this is the case, where does the escalation end? De-escalation is essentially a game of chicken – one side has to be content with not responding to an attack to take the temperature down.

My ECONS101 class has been covering game theory this week, including the chicken game. In the traditional game of chicken there are two rivals in cars, one at each end of the same street. They drive towards each other at top speed, and whichever rival swerves away first loses the game. So, each rival can choose to speed ahead or swerve away, and each would prefer to speed ahead and win the game. However, the problem is that if both simply keep speeding ahead, it will end in a disastrous crash.

I also recently read the book Hidden Games, by Moshe Hoffman and Erez Yoeli (which I reviewed here). Hoffman and Yoeli have an interesting section in the book on the hawk-dove game, which essentially the chicken game but with a slightly different motivating context. In the hawk-dove game, two rivals are competing over some resource. Each rival can choose to be aggressive or submissive, and whichever rival is more aggressive will win the resource. Each rival would prefer to be aggressive and win the resource. However, if both are aggressive, it ends in a massively disastrous battle.

Coming back to the case of Iran and Israel, this is clearly an example of the hawk-dove game (or the chicken game, if you prefer). This game is laid out in the payoff table below, where the strategies for Israel and Iran are to be aggressive, or submissive. The payoffs are expressed as "+" for good outcomes, and "-" for bad outcomes (and "--" is particularly bad), while zero is a neutral payoff.

To find the Nash equilibrium in this game, we use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium). In this game, the best responses are:

  1. If Iran chooses to be aggressive, Israel's best response is to be submissive (since "-" is better than "--" as a payoff - in other words, taking a bit of punishment is better than a massively disastrous war) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If Iran chooses to be submissive, Israel's best response is to be aggressive (since "+" is better than "0" as a payoff);
  3. If Israel chooses to be aggressive, Iran's best response is to be submissive (since "-" is better than "--" as a payoff); and
  4. If Israel chooses to be submissive, Iran's best response is to be aggressive (since "+" is better than "0" as a payoff).

In this scenario, there are no dominant strategies. Neither country has a strategy that is always better for them, no matter what the other country chooses to do. However, there are two Nash equilibriums (outcomes where both players are playing their best response), which occur when one country is aggressive, and the other is submissive.

The thing about the Iran-Israel hawk-dove game is that it isn't really a simultaneous game, as shown in the table above. It is a sequential game. Each player chooses whether to be aggressive or submissive, knowing what the other player chose to do previously. That sequential game is shown below. [*]

We can solve a sequential game using 'backward induction', which is essentially the same as the best response method, except we make sure we start with the last player, and work our way backwards through the game to work out what the first player should do. The resulting equilibrium that we find will be a 'subgame perfect Nash equilibrium'. In this case:

  1. If Israel chooses to be aggressive, Iran's best response is to be submissive (since "-" is better than "--" as a payoff) - now, since Iran would never choose to be aggressive when Israel has already been aggressive, Israel knows that the outcome will be that Iran is submissive;
  2. If Israel chooses to be submissive, Iran's best response is to be aggressive (since "+" is better than "0" as a payoff) - now, since Iran would never choose to be submissive when Israel has already been submissive, Israel knows that the outcome will be that Iran is aggressive;
  3. Israel will choose to be aggressive (since "+" is better than "-" as a payoff).

The subgame perfect Nash equilibrium is that Israel is aggressive, and Iran is submissive. As it turns out, that's sort of what happened. After an initial flurry of missile attacks, Iran stopped escalating the conflict.

One last thing to note is that this is a repeated game. Israel and Iran will find themselves in conflict often. The games as outlined above suggest that whichever country is the initial aggressor will end up getting their way, because the country moving second in the sequential game will be better off being submissive than retaliating. However, in repeated games, the outcome can often deviate from the equilibrium for strategic reasons.

Israel clearly doesn't want Iran to be aggressive, even though aggression would be good for Iran, if they moved first in this game. So, Israel wants to convince Iran not to make an aggressive first move. The only way that Israel can do that is to convince Iran that Iran would be worse off by making an aggressive first move. Israel needs to convince Iran that Israel will always retaliate with aggression. That would deter Iran from being aggressive as a first move. How does Israel achieve this? By developing a reputation for aggressively retaliating against any aggression. And indeed, that is what Israel has done (one need only look at Gaza or Lebanon for confirmation of this).

Israel's aggressive response to attacks by Iran, Gaza, and Lebanon is part of a strategic plan to deter future aggression against Israel. Many of us may not like it, but it's strategically rational. Whether it has a lasting effect remains to be seen.

*****

[*] I'm showing the game as having Israel move first. However, if you read Thomas's article, you'll see that 'who started it' is actually contested. I'm not taking a stand on that here, and in fact the game looks identical if Iran moves first.

Tuesday, 21 January 2025

Book review: Hidden games

Game theory has a lot of real-world applications. I am never short of good examples to use when teaching game theory in my ECONS101 class. However, I can always use more examples. And so, I was really interested to read Hidden Games, by Moshe Hoffman and Erez Yoeli. The subtitle promises: "The surprising power of game theory to explain irrational human behavior". I set aside the word 'surprising', as I wasn't expecting to be surprised, but I was expecting to be entertained.

And indeed, the book is entertaining. Hoffman and Yoeli use examples from several television series, including The Wire, The Sopranos, and even Star Trek: The Next Generation. I really enjoyed those examples, and many others. At one point, they use game theory to explain Elizabeth's decision about whether to marry Mr Darcy or not, in Jane Austen's Pride and Prejudice.

The overall aim of the book is to explain social puzzles, and Hoffman and Yoeli note that:

To explain all of our social puzzles, we will use game theory, but the game theory will often be hidden and will need to be interpreted through the lens of learning and evolutionary processes.

Moreover, they write that:

One of the premises on which the analyses in this book rest [sic] is that learning, regardless of whether it is from one's own experience via reinforcement or from others' via imitation and instruction, leads us to do what is good for us, at least on average, much of the time.

So, the overall theme of the book is that game theory can explain social puzzles, and that we humans (as well as other animals in some sections of the book) act as if we are solving these puzzles using game theory, and that is because of learning. Given that some of the learning is social learning, this is really an evolutionary argument. And it makes a good story.

The first few chapters are easy to read and follow, and will engage most readers. Even the maths (and game theory can have a lot of maths) is relatively straightforward, and in any case is explained in a way that is easy to understand. However, this takes a turn when the book gets to Chapter 8 where Hoffman and Yoeli introduce elements of Bayesian reasoning (and evidence) into the picture. The maths transitions towards greater difficulty, and the explanations are not as clear. From that point on, I understood the maths but still found the book to be heavy going. Overall, when Hoffman and Yoeli are using narrative examples, the book is good. When they resort to maths, which is all too frequent through the second half of the book, it is not so good. In fact, I think that the book would have been much better if the maths had been excised and the examples explained narratively without the complicated technical details.

I also found several parts where I thought a bit more depth of narrative would have helped. For example, Chapter 7 discusses 'countersignalling' (signalling a positive attribute by not signalling that one has that attribute). However, Hoffman and Yoeli don't explain how it is that someone without the positive attribute couldn't simply pretend to be countersignalling. 

This is a good, if somewhat uneven book. A game theory enthusiast would certainly enjoy it, as will the more maths-inclined reader. Those without a good understanding of maths will probably be turned off by that aspect of the second half of the book which, although understandable, is a bit of a shame.

Thursday, 2 January 2025

Book review: Games Businesses Play

Given the centrality of game theory to an understanding of business strategy, it seems surprising to me that the incorporation of game theory into IO (the economics field of industrial organisation) is a relatively recent phenomenon. That was one of the things that I learned reading Pankaj Ghemawat's 1997 book Games Businesses Play. Ghemawat is now a professor at the Stern School of Business at New York University, but at the time the book was written he was a professor at Harvard Business School. And so, as you might expect, the book employs a case study framework to the use of game theory in business situations. However, the scope of the book (and case studies) is necessarily narrowed, to examples of 'commitment decisions'. You can think of commitment decisions as a decision to commit to a particular course of action. For example, choosing whether to open a new factory, expand production, or enter a new market, are all commitment decisions, as are the reverse (closing a factory, decreasing production, or existing from an existing market). Ghemawat explains this choice of scope as follows:

The focus on commitment decisions as the unit of analysis is only slightly less obvious... game-theoretic IO might reasonably be expected to have the most of contribute to business strategy in regard to such decisions.

The book includes six main case studies, ranging from short-run to long-run decisions, and decisions in both expanding and declining industries. To some extent though, the cases themselves are less interesting than the general approach, given the dated nature of the material. However, the cases are also quite technical (and in particular, quite mathematical). This is not a book for a casual reader looking to pick up some simple tips on business strategy derived from game theoretic insights. That Harvard Business School students might cope with these cases really demonstrates the gulf in ability between those students and the 'average' business school (even the average graduate business school) student.

Ghemawat sees the book and the case studies as providing value in particular to those studying or researching in strategic management:

Taken together, the detailed case studies in this book strongly suggest that researchers in strategic management should increase their (currently low) level of attention to game theory instead of simply focusing on differential efficiency.

Sadly, I have seen little evidence that this advice has been heeded (and I say this as a failed strategic management undergraduate student from many years ago, as well as having observed students in our School's case competition over many years, who often seem to give little consideration to game theory in devising strategies for their given case).

I learned a lot from the book, but I suspect most readers won't appreciate the detailed technical approach. Unless you have a strong mathematical background (or a willingness to skip over the mathematical models and hope to pick up the story again afterwards), I would suggest this is not the book for you.

Sunday, 24 March 2024

In The Three Body Problem trilogy, Wallfacer Rey Diaz needed to better understand game theory

Regular readers of this blog may have noticed that I haven't posted a book review in a while. That's because I've been reading Cixin Liu's Three Body Problem trilogy (technically, the Remembrance of Earth's Past trilogy, an adaptation of which has just been released on Netflix as The Three Body Problem). I'm currently reading the second book, The Dark Forest.

Warning: Spoiler alert!

To give you some context, in the first book of the trilogy, Earth made contact with an alien civilisation, the Trisolarans. The Trisolaran fleet is currently on its way to Earth, in order to conquer us. Their homeworld is about to be destroyed, and their only hope of survival is to take over another planet. The fleet will take some 400 years to arrive, so Earth has some time to prepare. However, the Trisolarans have advanced technology, including deploying sophons, which are able to prevent Earth from conducting basic research in physics and other areas. So, Earth is stuck in a low-technology state, awaiting the arrival of the Trisolaran fleet. Even worse, the sophons can watch anything that happens on Earth and relay the information back to the Trisolarans, so Earth's preparations will be known to the Trisolaran fleet. To combat this, in the second book, Earth appoints four 'Wallfacers', who are given access to almost unlimited resources to execute plans that are known only to themselves, hidden from the rest of the Earth's population (and to the sophons, because the sophons can't read minds).

The second book of the trilogy is devoted to the Wallfacers and their plans (admittedly, I haven't finished reading it yet). I want to focus on the plans of Wallfacer Rey Diaz, whose plan involved planting large solar hydrogen bombs on Mercury, which when detonated would set off a chain reaction, destroying most of the solar system, including Earth. Diaz's plan was to negotiate with the Trisolaran fleet, warning them that if they didn't divert, Earth would be destroyed, sealing the fates of both the human and Trisolaran populations.

However, Wallfacer Rey Diaz's strategy is flawed. He needs to understand some basic game theory. To see why, consider the game shown in the payoff table below. The two players are Earth and the Trisolaran fleet (we'll assume that Diaz would choose strategy on behalf of Earth). Earth's two strategies are to blow up Mercury (detonate) or not. The Trisolaran fleet's two strategies are to continue to Earth, or divert. If Earth blows up Mercury, then Earth becomes extinct, regardless of what the Trisolaran fleet does. If Earth doesn't detonate Mercury, then Earth loses if the Trisolaran fleet continues, and wins if the Trisolaran fleet diverts. If the Trisolaran fleet diverts, they become extinct. If they continue to Earth, they become extinct if Earth blows up Mercury, but win if Earth does not.

To find the Nash equilibrium in this game, we use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium). In this game, the best responses are:

  1. If the Trisolaran fleet continues to Earth, Earth's best response is to not detonate (since losing is a better payoff than extinction [*]) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If the Trisolaran fleet diverts, Earth's best response is to not detonate (since winning is a better payoff than extinction);
  3. If Earth chooses to detonate, the Trisolaran fleet's best response is either option (since both payoffs are the same - extinction - both are best responses); and
  4. If Earth chooses not to detonate, the Trisolaran fleet's best response is to continue to Earth (since winning is a better payoff than extinction).

Note that Earth's best response is always to choose not to detonate. This is their dominant strategy. A player would always choose to play their dominant strategy, because choosing the other strategy makes them unambiguously worse off. And the Trisolarans would know this. This is what Wallfacer Rey Diaz gets wrong in his strategy. Earth won't blow Mercury up, and the Trisolarans know this, so there is no leverage for Earth in the negotiations.

The Trisolaran fleet has a weakly dominant strategy. Notice that continuing to Earth is always the Trisolaran fleet's best response. However, diverting is a best response if Earth chooses to detonate. So, continuing to Earth is not always better for the Trisolaran fleet, but it is never worse than the other strategy.

The single Nash equilibrium occurs where both players are playing a best response (where there are two ticks), which is where all Earth chooses not to detonate, and the Trisolaran fleet continues to Earth. It is little wonder then, that when Rey Diaz's strategy was revealed, the Earth governments were not happy. Not only was his strategy imperilling the Earth to the same extent as the Trisolarans, it was a strategy that simple game theory shows would not have succeeded.

*****

[*] You may wonder what the difference between losing and extinction is. Earth could lose, but some humans remain alive as slaves, or otherwise escape the planet before the Trisolarans arrive. It's not a great outcome, but better than extinction.

Sunday, 14 January 2024

Angola plays its dominant strategy in defecting against OPEC

The Financial Times reported last week (paywalled):

Angola, Africa’s second biggest oil producer, has said it is leaving Opec after disagreements over its production targets, delivering a blow to the oil cartel chaired by Saudi Arabia.

The decision comes after the producer group lowered Angola’s oil output target last month as part of a series of cuts led by Saudi Arabia to help prop up prices.

OPEC is an example of a cartel. Cartels can arise when a market is an oligopoly - a market where there are many buyers, but few sellers. A cartel essentially acts like a monopoly seller - it is able to use market power to extract greater economic rent from the market (in the form of higher profits, arising from higher prices), than the countries would be able to extract if they were competing with each other. The cartel can be maintained because there are few sellers, so it is relatively easy for them to coordinate their actions. In this case, OPEC coordinates to raise prices by restricting production.

However, there is always an incentive for each cartel member to cheat on the cartel agreement, or to leave the cartel entirely (as Angola has done). To see why, we can apply some game theory. Let's say that there are two players - Angola and 'the rest of OPEC'. Each player has two strategies - high production (which leads to lower prices and lower profits for oil producers), or low production (which leads to higher prices and higher profits). If one player has high production and the other low production, the high production player benefits more. However, if both players have high production, both are worse off. These outcomes and payoffs are illustrated in the diagram below (the payoff numbers represent profits, but are just made up to illustrate this example).

To find the Nash equilibrium in this game, we use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium). In this game, the best responses are:

  1. If the rest of OPEC chooses high production, Angola's best response is to choose high production (since 2 is a better payoff than 0) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If the rest of OPEC chooses low production, Angola's best response is to choose high production (since 4 is a better payoff than 3);
  3. If Angola chooses high production, the rest of OPEC's best response is to choose high production (since 10 is a better payoff than 9); and
  4. If Angola chooses low production, the rest of OPEC's best response is to choose high production (since 15 is a better payoff than 12).
Note that Angola's best response is always to choose high production. This is their dominant strategy. Likewise, the rest of OPEC's best response is always to choose high production, which makes it their dominant strategy as well. The single Nash equilibrium occurs where both players are playing a best response (where there are two ticks), which is where all of OPEC (including Angola) chooses high production.

Notice that both players would be unambiguously better off if they chose low production. However, both will choose high production, which makes them both worse off. This is a prisoners' dilemma game (it's a dilemma because, when both players act in their own best interests, both are made worse off).

That's not the end of this story though, because the simple example above assumes that this is a non-repeated game. A non-repeated game is played once only, after which the two players go their separate ways, never to interact again. Most games in the real world are not like that - they are repeated games. In a repeated game, the outcome may differ from the equilibrium of the non-repeated game, because the players can learn to work together to obtain the best outcome.

And that is what happens when a cartel forms. If all of OPEC (including Angola) works together and agrees to choose low production, both players benefit. That is what they were doing, up until Angola chose to leave OPEC. The problem here is that both players choosing low production is not an equilibrium. If Angola knows that the rest of OPEC is choosing low production, it is better off defecting from the agreement and choosing high production. Angola profits more that way (at least, in the short term).

So essentially, by leaving OPEC (and thereby choosing high production), Angola is simply playing its dominant strategy.

Read more:

Tuesday, 1 August 2023

The Tour de France, public goods, and the chicken game

I finally finished watching this year's Tour de France on Sunday. Yes, I was a week behind. That's because I was overseas when it started, and it took me that long to catch up (with big thanks to Sky On Demand!). Jonas Vingegaard well deserved his win. The individual time trial he rode on Stage 16 was amazing to watch (even if his team Jumbo Visma says so themselves).

Anyway, this is a blog about economics. Sports provide lots of great examples of economics in action, because economics is ultimately about choices, and so are sports. One striking example of economics in action in cycling road races occurs when there is a breakaway, and it is getting close to the finish line. The riders in the breakaway face a difficult choice. They can ride hard at the front of the breakaway, ensuring that the breakaway won't be caught by the peloton, and one of the breakaway riders will surely win the race. Or they can hold back, riding in the slipstream of the rider who is riding at the front, which lets them conserve energy for a sprint finish, but at the risk that the peloton catches them.

This exact scenario played out in Stage 18 of the Tour de France this year, with three riders approaching the finish. Victor Campenaerts rode hard towards the finish, ensuring the breakaway would succeed. However, it was Kasper Asgreen who won the stage, having conserved his energy for the final sprint among the breakaway riders.

Let's think about the incentives for a breakaway rider. Riding hard is a public good. It is non-rival (one cyclist benefiting from a rider riding hard at the front of the breakaway doesn't reduce the amount of the benefit available for the other riders in the breakaway) and non-excludable (if a rider is riding hard at the front of the breakaway, they can't easily prevent the other breakaway riders from sitting in their slipstream and conserving their energy).

Public goods, like riding hard at the front of the breakaway group, suffer from a free rider problem (pun intended!). Other riders can benefit from the front rider's hard work, without paying any of the cost themselves. It is difficult for a rider to justify riding hard at the front if other riders are unwilling to contribute, since they face all of the cost of riding hard, but the benefit (in terms of a better chance of winning the race) goes to the other riders (the free riders).

Ordinarily, the provision of public goods breaks down. They cannot be privately provided, because of the free rider problem. In this case though, cycling has developed norms that ensure some cooperation within the breakaway group. The riders tend to take turns at the front of the breakaway group, helping to increase the chances of success. However, the closer the race gets to the finish, the greater the incentives to free ride become. Regular cycling fans will no doubt remember many instances where a breakaway group has been caught, within sight of the finish line, because they failed to work together.

Another way of thinking about the incentives within a breakaway group is to use game theory. To make the problem simpler, let's say that the breakaway group only consists of two riders, and there are two strategies: (1) to ride hard; or (2) to hold back. We'll assume each rider makes their decision just once, and they make their decisions at the same time (a simultaneous game).  The payoffs for this scenario are shown in the table below. If both riders ride hard, they have a 50% chance of winning the race (since they will both be equally tired). If one rider rides hard and the other holds back, the rider that holds back wins the race for sure. If both riders hold back, then they are caught by the peloton, and neither of them wins (and they don't even finish in the top two in the race). What will happen?

To find the Nash equilibrium in this game, we use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium). In this game, the best responses are:

  1. If Rider B chooses to ride hard, Rider A's best response is to hold back (since winning for sure is better than a 50/50 chance of winning) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If Rider B chooses to hold back, Rider A's best response is to ride hard (since losing and finishing in the top two is better being caught by the peloton and finishing much lower in the order);
  3. If Rider A chooses to ride hard, Rider B's best response is to hold back (since winning for sure is better than a 50/50 chance of winning); and
  4. If Rider A chooses to hold back, Rider B's best response is to ride hard (since losing and finishing in the top two is better being caught by the peloton and finishing much lower in the order).

In this scenario, there are no dominant strategies. Neither rider has a strategy that is always better for them, no matter what the other rider chooses to do. However, there are two Nash equilibriums (outcomes where both players are playing their best response), which occur when one rider rides hard, and the other holds back. Neither rider will want to be the rider that rides hard, so both may be holding out hoping that the other rider will ride hard. This is the free rider problem described earlier. This game is an example of the chicken game (which I have discussed here). If both riders hold back, hoping that the other rider will ride hard, both riders will be caught by the peloton.

The chicken game is an example of a coordination game. To end up at one of the equilibriums (or another), the players need to coordinate their actions. However, in this case neither rider really wants to coordinate on the other rider's preferred equilibrium. Both really want to hold back, especially closer to the finish line, which is why the breakaway can often be caught.

Riders are motivated by the chance to win the race. That is why breakaway groups form in the first place. However, the incentives outlined above work against the breakaway succeeding. And riders are aware of these issues. One thing that often happens is that, towards the end of a race, one rider will ride especially hard, breaking away from the breakaway group. There is no free rider problem when a rider is riding by themselves. Sadly, solo breakaways are seldom successful (except on mountain stages), because the effort required to remain clear from a group of breakaway riders who suddenly become more motivated to work together and catch the solo breakaway rider is very high. The solo breakaway rider is often caught, after which the chicken game and free riding begins again.

One thing that can increase the success of a breakaway is to have multiple teammates in the breakaway group. Teammates are more likely (but not certain) to be able to coordinate their strategies, and work together, reducing the free riding problem. That's why riders in the peloton are more vigilant and energetic in chasing down an early breakaway group that has multiple riders from the same team. Most of the time, a breakaway group will only go clear if every rider in the group is from a different team. Riders in the peloton don't want the breakaway to succeed, and having all breakaway riders from different teams decreases the chance that a rider from the breakaway wins the race.

There is a lot of strategy in sports, and cycling is no exception. There are also a lot of choices for athletes to make, and choices involves trade-offs. That, along with the transparent rules and the obvious goals of the athletes involved (they want to win), is why sports can provide a lot of useful illustrations of economic concepts.

Monday, 10 April 2023

The game theory of an AI pause

My news feed has been dominated over the last week by arguments both for and against a pause on AI development, prompted by this open letter by the Future of Life Institute., which called for:

AI labs to immediately pause for at least 6 months the training of AI systems more powerful than GPT-4.

Chas Hasbrouck has an excellent post summarising the various views on AI held by different groups (and examples of the people belonging to each group). Tyler Cowen then suggested on the Marginal Revolution blog that we should be considering the game theory of this situation (see also his column on Bloomberg - as Hasbrouck notes, Cowen is one of the 'Pragmatists'). I want to follow up Cowen's suggestion, and look at the game theory. However, things are complicated a little, because it isn't clear what the payoffs are in this game. There is so much uncertainty. So, in this post, I present three different scenarios, and work through the game theory of each of them. For simplicity, each game has two players (call them Country A and Country B), and each player has two strategies (pause development on AI, or speed ahead).

Scenario #1: AI Doom with any development

In this scenario, if either country speeds ahead and the other doesn't, the outcomes are bad, but if both countries speed ahead, the planet faces an extinction-level event (for humans, at the least). The payoffs for this scenario are shown in the table below.

To find the Nash equilibrium in this game, we use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium). In this game, the best responses are:

  1. If Country B chooses to pause development, Country A's best response is to pause development (since a payoff of 0 is better than a payoff of -5) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If Country B chooses to speed ahead, Country A's best response is to pause development (since a payoff of -10 is better than extinction);
  3. If Country A chooses to pause development, Country B's best response is to pause development (since a payoff of 0 is better than a payoff of -5); and
  4. If Country A chooses to speed ahead, Country B's best response is to pause development (since a payoff of -10 is better than extinction).

In this scenario, both countries have a dominant strategy to pause development. Pausing development is always better for a country, no matter what the other country decides to do (pausing development is always the best response).

For anyone who believes in this scenario, pausing development will seem like a no-brainer, since it is a dominant strategy.

Scenario #2: AI Doom if everyone speeds ahead

In this scenario, if both countries speed ahead, the planet faces an extinction-level event (for humans, at the least). However, if only one country speeds ahead, then AI alignment can keep up, preventing the extinction-level event. The country that speeds ahead earns a big advantage. The payoffs for this scenario are shown in the table below.

Again, let's find the Nash equilibrium using the best response method. In this game, the best responses are:

  1. If Country B chooses to pause development, Country A's best response is to speed ahead (since a payoff of 10 is better than a payoff of 0);
  2. If Country B chooses to speed ahead, Country A's best response is to pause development (since a payoff of -2 is better than extinction);
  3. If Country A chooses to pause development, Country B's best response is to speed ahead (since a payoff of 10 is better than a payoff of 0); and
  4. If Country A chooses to speed ahead, Country B's best response is to pause development (since a payoff of -2 is better than extinction).

In this scenario, there is no dominant strategy. However, there are two Nash equilibriums, which occur when one country speeds ahead, and the other pauses development. Neither country will want to be the country that pauses, so both will be holding out hoping that the other country will pause. This is an example of the chicken game (which I have discussed here). If both countries speed ahead, hoping that the other country will pause, we will end up with an extinction-level event.

For anyone who believes in this scenario, pausing development will seem like a good option, even if only one country will pause development. However, no country is going to want to willingly buy into pausing development.

Scenario #3: AI Utopia

In this scenario, if both countries speed ahead, the planet reaches an AI utopia. The fears of an extinction-level event do not play out, and everyone is gloriously happy. However, if only one country speeds ahead, then the outcomes are good, but not as good as they would be if both countries sped ahead. Also, the country that speeds ahead earns a big advantage. The payoffs for this scenario are shown in the table below.

Again, let's find the Nash equilibrium using the best response method. In this game, the best responses are:

  1. If Country B chooses to pause development, Country A's best response is to speed ahead (since a payoff of 10 is better than a payoff of 0);
  2. If Country B chooses to speed ahead, Country A's best response is to speed ahead (since utopia is better than a payoff of -2);
  3. If Country A chooses to pause development, Country B's best response is to speed ahead (since a payoff of 10 is better than a payoff of 0); and
  4. If Country A chooses to speed ahead, Country B's best response is to speed ahead (since utopia is better than a payoff of -2).

In this scenario, both countries have a dominant strategy to speed ahead. Speeding ahead is always better for a country, no matter what the other country decides to do (speeding ahead is always the best response).

For anyone who believes in this scenario, speeding ahead will seem like a no-brainer, since it is a dominant strategy.

Which is the 'true' scenario? I have no idea. No one has any idea. We could ask ChatGPT, but I strongly suspect that ChatGPT will have no idea as well. [*] What the experts believe we should do depends on which of the scenarios they believe is likely to be playing out. Or perhaps, with a chance that any of the three scenarios (or any other of millions of other potential scenarios with different players and payoffs) is playing out, perhaps the precautionary principle should apply? The problem there, though, is if any country pauses development, the best response in any of the scenarios except the first one is for other countries to speed ahead. So, unless all countries can be convinced to apply the precautionary principle, pausing development is simply unlikely.

We live in interesting times.

*****

[*] Actually, I tried this, and ChatGPT refused to offer an opinion, instead it said: "...it is crucial that policymakers and stakeholders work together to develop standards and guidelines for responsible AI development and deployment to minimize potential risks and maximize benefits for society as a whole." Thanks ChatGPT.

Sunday, 2 April 2023

How not to strategise for penalty kicks

In game theory, a pure strategy is an unconditional choice of strategy for a player. In other words, the player chooses that strategy for sure. That distinguishes it from mixed strategy, where the player randomises their actions, choosing each of the possible strategies with some probability (which might be zero). There are lots of examples of mixed strategies. One that I use in my ECONS101 class is the choice for a tennis player over whether to serve down the middle, into the body, or out wide. If they chose one strategy for sure, they would reduce their chances of winning. Instead, they should randomise - sometimes choosing the first strategy, sometimes the second, and sometimes the third.

Another example from sports is the penalty kick in football (or soccer, if you prefer). The penalty taker must choose which side to kick towards, and the goalkeeper must choose which way to defend. I've discussed this game and the mixed strategy equilibrium before (see here and here).

The key problem with mixed strategy is that it genuinely involves randomisation. You cannot reason a pure strategy solution to a mixed strategy game. If you do, you end up with something like this:

I'm not sure where the video comes from (TV or movies, or something else), but it is very similar to a story related in the book Soccernomics, by Simon Kuper and Stefan Szymanski (as Robbie Butler notes here). The solution to mixed strategy games is not to try and solve them with pure strategy, but to randomise.

[HT: Jadrian Wooten at Critical Commons, via the Economics Media Library]

Read more: