Showing posts with label Mixed strategy. Show all posts
Showing posts with label Mixed strategy. Show all posts

Sunday, 2 April 2023

How not to strategise for penalty kicks

In game theory, a pure strategy is an unconditional choice of strategy for a player. In other words, the player chooses that strategy for sure. That distinguishes it from mixed strategy, where the player randomises their actions, choosing each of the possible strategies with some probability (which might be zero). There are lots of examples of mixed strategies. One that I use in my ECONS101 class is the choice for a tennis player over whether to serve down the middle, into the body, or out wide. If they chose one strategy for sure, they would reduce their chances of winning. Instead, they should randomise - sometimes choosing the first strategy, sometimes the second, and sometimes the third.

Another example from sports is the penalty kick in football (or soccer, if you prefer). The penalty taker must choose which side to kick towards, and the goalkeeper must choose which way to defend. I've discussed this game and the mixed strategy equilibrium before (see here and here).

The key problem with mixed strategy is that it genuinely involves randomisation. You cannot reason a pure strategy solution to a mixed strategy game. If you do, you end up with something like this:

I'm not sure where the video comes from (TV or movies, or something else), but it is very similar to a story related in the book Soccernomics, by Simon Kuper and Stefan Szymanski (as Robbie Butler notes here). The solution to mixed strategy games is not to try and solve them with pure strategy, but to randomise.

[HT: Jadrian Wooten at Critical Commons, via the Economics Media Library]

Read more:

Tuesday, 16 February 2021

Combating cheating in online tests

From my perspective, the most challenging aspect of teaching during the pandemic lockdowns last year wasn't the teaching itself, it was dealing with students cheating in the online assessment. To give you some idea, I sent more students to the Student Discipline Committee in B Trimester 2020 than I had in the previous 10 years of teaching combined. All but one of those students ended up failing their paper. And I was not alone. The Student Disciplinary Committee faced a huge increase in workload, especially related to students using contract cheating websites to answer assessment questions for them.

Anyway, as you may expect, my experiences (and those of my colleagues) are not isolated examples. In a new paper in the Journal of Economic Behavior and Organization (ungated earlier version here), Eren Bilen (University of South Carolina) and Alexander Matros (Lancaster University) looked at cheating in online assessments. They use two examples to illustrate the pervasiveness of cheating: (1) students in an intermediate level class in Spring Semester 2020 (when lockdowns were introduced partway through the semester); and (2) online chess tournaments. They motivate their analysis with a simple game theoretic model, as shown below (the first payoff is to the student, and the second payoff is to the professor).


They note that in the sequential game:

It is easy to find a unique subgame perfect equilibrium outcome, where the student is honest and the professor does not report the student. Note that this is the best outcome for the professor and the second best outcome for the student.

To see why that is the subgame perfect Nash equilibrium, we can use backward induction. Essentially, we work out what the second player (the professor) will do first, and then use that to work out what the first player (the student) will do. In this case, if the student cheats, then we are moving down the left branch of the tree. The best option for the professor in that case is to report the student (since a payoff of 3 is better than a payoff of 2). So, the student knows that if they cheat, the professor will report them. Now, if the student doesn't cheat, then we are moving down the right branch of the tree. The best option for the professor in that case is not to report the student (since a payoff of 4 is better than a payoff of 1). So, the student knows that if they don't cheat, the professor will not report them. So, the choice for the student is to cheat and get reported (and receive a payoff of 1) or not cheat and not get reported (and receive a payoff of 3). Of course, the student will choose not to cheat. The subgame perfect Nash equilibrium here is that the student doesn't cheat, and the professor doesn't report them.

The problem with that analysis is that the professor doesn't know with certainty if the student has cheated or not. So, Bilen and Matros move onto a sequential game, as shown below. Even though the players make their choices sequentially, because the student's choice about whether to cheat or not is not revealed to the professor, it is as if the professor is making their choice about whether to report or not at the same time as the student. That makes this a simultaneous game.



Bilen and Matros note that, in this game:
This game has a unique mixed-strategy equilibrium, which means that the student and the professor should randomize between their two actions in equilibrium. Thus cheating as well as reporting is a part of the equilibrium.
To see why, we need to try to find the Nash equilibriums in this game, and to do that we can use the 'best response method'. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the textbook definition of Nash equilibrium). In this game, the best responses are:
  1. If the student chooses to cheat, the professor's best response is to report the student (since 3 is a better payoff than 2);
  2. If the student chooses not to cheat, the professor's best response is not to report the student (since 4 is a better payoff than 1);
  3. If the professor chooses to report the student, the student's best response is to not cheat (since 2 is a better payoff than 1); and
  4. If the professor chooses not to report the student, the student's best response is to cheat (since 4 is a better payoff than 3).
A Nash equilibrium occurs where both players' best responses coincide (normally I would track this with ticks and crosses, but since I didn't create the payoff table I haven't done so in this case). Notice that there isn't actually any case where both players are playing a best response. If the student cheats, the professor's best response is to report them. But if the professor is going to report the student, the student's best response is to not cheat. But if the student doesn't cheat, the professor's best response is not to report them. But if the professor doesn't report the student, the student's best response is to cheat. We are simply going around in a circle.

In cases such as this, we say that there is no Nash equilibrium in pure strategy. However, there will be a mixed strategy equilibrium, where the players randomise their choices of strategy. The student should cheat with some probability, and the professor should report the student with some probability. The optimal probabilities depend on the actual payoffs to the players (we could work it out in this case, but I'm not going to go that far today).

Anyway, Bilen and Matros conclude that in-person tests and examinations most closely resemble the sequential game, where the equilibrium is for students not to cheat, but that online tests most closely resemble the simultaneous game, where cheating will be much more common. And, their analysis supports that theory:
Using a simple way to detect cheating - timestamps from the students’ Access Logs - we identify cases where students were able to type in their answers under thirty seconds per question. We found that the solution keys for the exam were distributed online, and these students typed in the correct as well as incorrect answers using the solution keys they had at hand.

Then they present their proposed solution to the problem of online cheating:

In order to address this issue based on our theoretical models, we suggest that instructors present their students with two options: (1) If a student voluntarily agrees to use a camera to record themselves while taking an exam, this record can be used as evidence of innocence if the student is accused of cheating; (2) If the student refuses to use a camera due to privacy concerns, the instructor should be allowed to make the final decision on whether or not the student is guilty of cheating, with evidence of cheating remaining private to the instructor.

I'm not sure that I agree. The optimal solution would be one that returns to conditions where cheating is easy to detect, as in the sequential game above. A voluntary webcam doesn't do this, since students who want to cheat, as well as students who have privacy concerns, would opt out. The game would revert to the simultaneous game for those students, and some of those students would cheat.

My solution, which will be implemented if we go into lockdown and can't have a final examination in the trimester due to start in a couple of weeks, is to move to individual oral examinations. It is a lot harder to hide when you are put on the spot in a Zoom call and asked to answer a question chosen at random by the lecturer. Certainly, you can't use an online cheating website to provide the answer for you in such a situation! I haven't fully worked out the mechanics of how it would work (and hopefully I never have to implement it!), but it would seem to me to return the assessment to a sequential game.

On the other hand, the theory of asymmetric information and signalling does suggest that the webcam solution may work. The student knows whether they intend to cheat or not. The professor doesn't know. Whether the student is planning to cheat is private information. The student can reveal this information by the choices they make. Say that the professor offers a voluntary webcam policy, where students who agree set up a webcam such that they can be observed while they complete the test. Students who agree to the webcam are clearly those that weren't intending to cheat, because then they would surely be caught. In contrast, if a student was intending to cheat, they wouldn't agree to the policy. And so, the professor is able to reveal who the likely cheaters are, simply by offering the policy. The students signal whether they are cheaters by agreeing to have the webcam, or not.

That seems way too simple, and you could argue that it is somewhat coercive, forcing honest students to compromise their privacy to signal their honesty. It's going to be imperfect too, because it would be difficult to separate students who chose not to webcam from those that can't afford a webcam, or whose internet connection is too slow to support webcam use, and so on.

And invigilating the webcam solution is going to be extraordinarily expensive. You can't get AI to do the supervision (at least, not yet). And since Zoom can only support 25 faces on screen at a time, you would need at least one invigilator per 25 students. Probably you need more than that, because the invigilator needs to be able to see clearly what the student is doing (so unless they are using a 50" screen, you probably want a whole lot fewer than 25 images on-screen at a time). I think I'll stick to the oral examination solution.

Cheating in online assessments is clearly a serious problem. This is the main reason that I am highly sceptical about the move to online education. Until we can solve the serious academic integrity issues in a cost-effective way (and we currently don't have one), the quality of assessment (in terms of accurately ranking students or assessing their knowledge and skills) is so flow that it makes a mockery of the whole idea of teaching online.

Monday, 20 February 2017

The irrationality of NFL play-callers

I recently read two papers that both essentially conclude (based on different aspects) that NFL play callers are not rational (or more specifically, not rational and risk neutral - an important point I'll return to at the end of the post). Recall that a rational decision-maker weighs up the costs and benefits of a decision, and when faced with mutually exclusive options (such as choosing which play to run in an NFL game), they should choose the option with the greatest net benefit (benefits minus costs).

The first paper (by Jonathan Hartley, an MBA student at the Wharton School at the University of Pennsylvania) looks at play-callers' choices between an extra point attempt and a two-point attempt following a touchdown. A rational and risk-neutral play-caller should choose whichever play provides the greatest expected benefit (expected number of points). In this case, Hartley found:
Between 2002 and 2014, the extra point conversion rate was 99.2% (out of 7738 attempts). As the average two point conversion rate remained 0.475, the expected points from a two-point conversion remains 0.95 below the automatic 0.992 points...
Over 2 seasons since the implementation of the new rules [increasing the distance the extra point try is attempted from], the extra point conversion rate has fallen from 0.992 to 0.95. Moreover, the total number of 2 point conversion attempts per season has nearly doubled...
In other words, when the NFL changed the extra point to being attempted from a greater distance (thereby making it more difficult), the expected value of an extra point try fell from 0.992 to 0.95 points. The expected value of a two-point conversion remained steady at 0.95 points. So, a rational and risk-neutral play-caller should now be indifferent between an extra point attempt and a two-point conversion. However, as Hartley shows in the paper, most teams still attempt very few two-point conversions, even those teams that have a history of success at them. The paper itself is pretty rough, but I wish the MBA students here could do this sort of work!

The second paper, by Noha Emara (Rutgers), David Owens (Haverford College), John Smith (Rutgers), and Lisa Wilmer (Florida State), is forthcoming in the Journal of Behavioral and Experimental Economics (ungated earlier version here), and looks at serial correlation in play-calling. Serial correlation occurs when you have a time series (like a series of plays) and where each observation in the time series is related (positively or negatively) to the observation or observations earlier in the time series. Obviously, an NFL offensive play-caller wants to call players in a random way - what we call a mixed strategy. Mixed strategy is particularly important in sports - think of the choice of where to serve in tennis, or where to shoot a penalty or which way to dive as a goalkeeper in soccer (see here or here for more on this). If an NFL offensive play-caller doesn't effectively randomise their play calling, then the defence can potentially exploit some prior knowledge of the play about to be called.

Humans are rubbish at trying to create random series, and indeed that's what Emara et al. found, based on their dataset of more than 200,000 plays from the 2000-2012 NFL seasons:
...the previous pass variable is negative and significant in each specification. This provides evidence that, even after controlling for down, distance, field position, and other observables, play calling exhibits significant negative serial correlation. The Previous pass-Previous failure interaction estimate is negative significant in both of the specifications where it appears, suggesting that play calling becomes even more negatively serially correlated following a failed play.
To translate, play-callers are significantly more likely to call a running play after a previous passing play, and to call a passing play after a previous running play, than would be expected if they were selecting plays randomly. And on top of that, if the previous play was a failure (e.g. if it lost yards), then they are even more likely to change the play type on the following play.

To make things worse, Emara et al. find evidence that teams would be better off if they ran more plays that were the same as the previous play:
We find that a rush following a rush gains 0.14 more yards than a rush following a pass. We also find that a pass following a pass gains 0.21 more yards than a pass following a rush. Estimates are more pronounced when we also control for whether the previous play was a failure. We find that a rush gains 0.24 more yards more following a failed rush than following a failed pass. Also, a pass gains 0.34 more yards following a non-failed pass than following a non-failed rush.
In summary, we find evidence that the efficacy of a play, as measured by yards gained, increases if it follows a play of the same type.
The results is even stronger on second down plays, but not so much for third down plays. However, the take-away message, like that of the first paper, is that play-callers are not being purely rational.

However, there is a caveat here. If we think that, based on this evidence, that play-callers should be calling more two-point conversions and switching up play types less often, then we may be forgetting that there is also a wider game at play here. If play-callers are risk averse, then this affects their decision-making. The two-point conversion may have the same expected value as an extra point attempt, but it is riskier (see also this post on NBA three-pointers from last week), so a risk averse play caller may avoid the two-point conversion more than the simple comparison of expected values would suggest.

But what about the play-callers in the second paper? Emara et al. have thought about this, and this is what they offer:
Perhaps teams feel pressure not to repeat the play type on offense, in order to avoid criticism for being too “predictable” by fans, media, or executives who have difficulty detecting whether outcomes of a sequence are statistically independent. Further, perhaps this concern is sufficiently important so that teams accept the negative consequences that arise from the risk that the defense can detect a pattern in their mixing.
Making play calls that the fans think are predictable (but which are actually more random) may make the play-caller themselves at risk of losing their job (or at least, of looking like they are doing a poor job). So, play-callers may attempt to make their play calls look more random by switching (from run to pass or vice versa) more often than they should, even though this is actually less random and costs the team in terms of yards gained per play.

The question is, now that these trends are known, will any team want to exploit them?

[HT: Marginal Revolution, here and here]

Monday, 9 January 2017

The game theory of withholding supply

Back in October, Ford stopped producing the Falcon XR8 Sprint. What happened? According to the Daily Telegraph:
FORD dealers are charging a staggering $30,000 more than the recommended retail price — up from $60,000 to $90,000 — for the final Falcon V8 sedans as buyers try to secure a future classic.
The last batch of Falcon V8s was thought to have sold out, but some dealers held a secret stash to release them onto the market in the final days of production so they could jack up the price.
Why did the price rise? Ford Australia boss Graeme Whickman explains:
“It’s supply and demand,” said Mr Whickman. “We set a wholesale price and recommended retail price … but at the end of the day the dealer and the customer decide what the vehicle is going to be sold and bought for,” he said, referring to high prices being charged for the initial shipment of Mustangs last year.
Demand for the Falcon XR8 was high, and the sellers withheld some supply - high demand and lower supply ensures that prices will increase. in this case by up to 50%.

So, why don't firms do this all the time? Why not withhold supply all the time? A bit of game theory can help, as laid out in the payoff table below [*]. The seller can choose to withhold stock, or not. The buyer can choose to buy now, or wait and buy later.


Where is the Nash equilibrium in this game? Consider the seller's choice first. If the buyers choose to buy now, the seller is better off choosing to withhold some stock, because profits will be higher. If the buyers choose to wait and buy later, the seller is better off choosing to not withhold stock, because profits will be higher. So, the seller doesn't have a dominant strategy (a strategy that is always better for them, no matter what the buyers choose to do).

Now consider the buyers' choice. If the seller chooses to withhold some stock, the buyers are better off choosing to wait and buy later, since they will force the price to fall when the withheld stock is released.  If the seller chooses not to withhold stock, the buyers are better off to buy now, or they will miss out on a car. So, the buyers don't have a dominant strategy either.

Neither player has a dominant strategy, so what is the solution to this game? We can find it using the best response method (which we've already described in the last two paragraphs). Any combination of strategies where both players are choosing their best response to the other player's strategy is a Nash equilibrium. However, in this case there is no Nash equilibrium (at least, no equilibrium in pure strategy). Instead, there will be a mixed strategy equilibrium where both the seller and the buyers should randomise their actions. The seller should sometimes withhold stock, and other times not.

Of course, that analysis assumes the game is not repeated. Car sellers sell new models each year, so really this is a repeated game. How does this change the game? If a game is repeated, then players can develop reputations. So, this actually reinforces the importance of the seller not constantly withholding stock. If they develop a reputation for withholding stock to sell later, then buyers will recognise that the seller's prices are artificially high and will wait until the seller releases the withheld stock, lowering the price.

So, provided buyers (as a group) learn from this experience, then there is little to fear from Ford dealers repeating the exercise every time a popular car is withdrawn from production. The problem is, of course, that it is unlikely to be the same group of buyers next time around.

*****

[*] This game is presented as a simultaneous game, even though the seller clearly chooses their strategy before the buyer. This is because when the buyer chooses to buy now or wait, they don't know whether the seller has withheld stock or not. For simplicity, we're also assuming that all buyers act the same way, when of course they won't. However, the overall point still stands even if we have multiple buyers in the game.

Saturday, 20 June 2015

More on game theory and the penalty shootout

A couple of weeks ago I wrote a post on game theory and penalty shootouts:
The penalty shootout is an excellent example of a game with no Nash equilibrium in pure strategy - where the equilibrium is a mixed strategy equilibrium. That is, it is best for the players to randomise their strategy choice...
Notice that there are no outcomes (cells) where both players are playing a best response, so there is no Nash equilibrium in pure strategy. To work out the mixed strategy equilibrium, we would need to know a bit more about the probability of success if both goalkeeper and shooter chose the same (here we assume there was a 0% chance of a goal), and the probability of success if they chose differently (here we assumed there was a 100% chance of a goal). 
Mark Johnston (HOD economics at King's College in Auckland) pointed me to several blog posts he has written on penalty shootouts. This post in particular gives us the probability of success or failure depending on which way the kicker shoots and which way the goalkeeper dives (the data comes from Ignacio Palacios-Huerta's book Beautiful Game Theory):


In the table above, NS represents the kicker's 'natural side' (to the right for a right-footed kicker), and OS represents the kicker's opposite side. Solving for the mixed strategy equilibrium requires some algebra. Let p be the probability that the kicker kicks to their natural side (and 1-p will be the probability that the kicker kicks to the opposite side). And let q be the probability that the goalkeepers dives to the kicker's natural side (and 1-q will be the probability that the goalkeeper dives to the kicker's opposite side).

Now, for each player we set the expected values of their two strategy choices to be equal, but those expected values depend on the probability that the other player chooses each strategy (this is analagous to our definition of Nash equilibrium, where both players are doing the best they can, given the choice of the other player).

So, for the kicker:
EV[natural side] = 0.7q + 0.95(1-q) = 0.95 - 0.25q
EV[opposite side] = 0.92q + 0.58(1-q) = 0.58 + 0.34q
0.95 - 0.25q = 0.58 + 0.34q
0.59q = 0.37
q = 0.627

And, for the goalkeeper (remembering that the payoffs to the goalkeeper are the complement of the payoffs to the kicker):
EV[natural side] = 0.3p + 0.08(1-p) = 0.08 + 0.22p
EV[opposite side] = 0.05p + 0.42(1-p) = 0.42 - 0.37p
0.08 + 0.22p = 0.42 - 0.37p
0.59p = 0.34
p = 0.576

So, the mixed strategy equilibrium is for the kicker to kick to their natural side 58% of the time (p = 0.576) and the goalkeeper to dive to the kicker's natural side 63% of the time (q = 0.627). Which makes sense - the kicker has a higher probability of scoring on their natural side, and knowing this the goalkeeper should go to that side more often.

For more on solving for mixed strategy equilibrium, try this video by William Spaniel:


Sunday, 7 June 2015

Game theory and the penalty shootout

The week before last in ECON100 we covered game theory. New Zealand is currently hosting the FIFA U20 World Cup, and Christoph Schumacher and Nigel Espie from Massey University have been writing a series of articles in the New Zealand Herald, linking the World Cup to business. Last week they wrote about penalty shootouts:
Seen through a game theory lens, a penalty shootout is a game that involves two players - the striker and the keeper. To simplify matters, let's assume that both players have three options. The striker can choose to make their shot to the right, centre, or left of the goal. Similarly, the keeper can elect to dive to the left or right, or defend the centre of the goal.
Let's also assume the striker is equally good at taking shots to all three areas in the goal and the keeper is equally good at saving balls kicked to the three sections of the goal. So what would be the best strategy for a goalkeeper in this situation?
The penalty shootout is an excellent example of a game with no Nash equilibrium in pure strategy - where the equilibrium is a mixed strategy equilibrium. That is, it is best for the players to randomise their strategy choice. As Schumacher and Espie put it:
With no past knowledge of the striker's penalty history, you would expect the striker to randomly select their shot. Game theory would suggest that the best option for the keeper is to also randomly select an area of the goal to defend. This maximises the likelihood of saving a goal while preventing the opposing team from discerning any preference by the keeper.
To see why, let's set up the game as we would in ECON100. It's a simultaneous game, because even though the shooter makes their choice first, it's unlikely that the goalkeeper has enough time to properly observe the shooter's choice before choosing their own strategy. So, we can represent the game in a payoff table (see below). The shooter has three strategy options (left, centre, right), and the goalkeeper has three strategy options (left, centre, right). To keep things simple, let's assume that if both players choose the same strategy then the goalkeeper saves the shot (the goalkeeper receives a payoff of +1, and the shooter receives a payoff of -1), and if both players choose different strategies, the goal is scored (the goalkeeper receives a payoff of -1, and the shooter receives a payoff of +1). The payoff tables looks as follows:


To confirm that there are no pure strategy Nash equilibriums, we can use the 'best response' method. To do this, we track: for each player, for each strategy, what is the best response of the other player. Where both players are selecting a best response, they are doing the best they can, given the choice of the other player (this is the definition of Nash equilibrium).

For our game outlined above:

  1. If the shooter shoots to the left, the goalkeeper's best response is to dive to the left (since +1 is better than -1) [we track the best responses with ticks, and not-best-responses with crosses; Note: I'm also tracking which payoffs I am comparing with numbers corresponding to the numbers in this list];
  2. If the shooter shoots to the centre, the goalkeeper's best response is to remain in the centre (since +1 is better than -1);
  3. If the shooter shoots to the right, the goalkeeper's best response is to dive to the right (since +1 is better than -1);
  4. If the goalkeeper dives to the left, the shooter's best response is to shoot to the centre or to the right (since +1 is better than -1);
  5. If the goalkeeper remains in the centre, the shooter's best response is to shoot to the left or to the right (since +1 is better than -1); and
  6. If the goalkeeper dives to the right, the shooter's best response is to shoot to the left or to the centre (since +1 is better than -1).
Notice that there are no outcomes (cells) where both players are playing a best response, so there is no Nash equilibrium in pure strategy. To work out the mixed strategy equilibrium, we would need to know a bit more about the probability of success if both goalkeeper and shooter chose the same (here we assume there was a 0% chance of a goal), and the probability of success if they chose differently (here we assumed there was a 100% chance of a goal). Moreover, these probabilities might be different depending on whether it was to the left or right. The equilibrium would therefore not necessarily be random and symmetric.


Sports are filled with mixed strategy equilibriums. Think about tennis serving (down the centre, into the body, out wide), or American football offense (pass vs. run), to name just two. In these (and many other cases), it is best to randomise your actions, because if you become too predictable, your rival can take advantage of that.