Monday, 11 April 2022

Online elite chess and cognitive performance during the pandemic

Does remote working increase productivity, or decrease productivity? The pandemic forced a lot of workers into remote working, so perhaps this natural experiment can give us some idea of the impacts of remote working. Do we gain more from avoiding commuting time, greater flexibility over work time and workspace, and fewer interruptions from colleagues, than we lose from reduced interaction, supervision and structure (in addition to whatever other effects might happen in either direction)? Despite the hype, the results so far are far from clear, especially in terms of what types of jobs or work improve in a remote setting.

An interesting new article by Steffen Künn, Christian Seel (both Maastricht University), and Dainis Zegners (Rotterdam School of Management), published in The Economic Journal (open access) provides a contribution towards answering those questions. Künn et al. look at the impact of the shift to online of elite chess tournaments. Specifically:

Our data consist of games from the World Rapid Chess Championships 2018–2019, played offline in Saint Petersburg and Moscow, and from the Magnus Carlsen Chess Tour and its sequel, the Champions Chess Tour, both played online from April to November 2020 on the internet chess platform chess24.com... the majority of players (20 out of 28) in the online tournaments also competed in at least one of the World Rapid Chess Championships in the years 2018–2019, enabling us to make within-player comparisons of performance for each of these 20 players.

Künn et al. measure the performance of each chess player for every move in every one of those tournaments (with a few exceptions, and excluding the first fifteen moves for each player in each game), relative to one of the top chess engines. As they explain:

To estimate the effect of playing online on chess players’ performance, we evaluate each move in each game in our sample using the chess engine Stockfish 11... 

For a given position in game g before individual move mig, the chess engine computes an evaluation of the position in terms of the pawn metric Pigm... The numerical value of the pawn metric indicates the size of the advantage from the perspective of player i, with one unit indicating an advantage that is comparable to being one pawn up...

Künn et al. use this evaluation to generate a measure of 'raw error', being the difference in the pawn metric between the player's choice of move, and the 'optimal' move as determined by Stockfish. They then compare this raw error between play in online tournaments and play in face-to-face tournaments, for the same players. They find that:

...playing online leads to a reduction in the quality of moves. The error variable... is, on average, 1.7 units larger when playing online than when playing identical moves in an offline setting. This corresponds to a 1.7% increase of the measure... or an approximately 7.5% increase in the RawError... The effect is statistically significant at the 5% level.

The effect is quite sizeable:

Playing online increases the error variable, on average, by 1.7 units, which corresponds to a loss of 130 points of Elo rating.

In reading the paper, my first thought was that the results would be contaminated by the psychological effects of the pandemic. Fear or anxiety could easily lead to suboptimal performance, and cause the observed increase in error, rather than reduced performance in the online format per se. However, Künn et al. anticipate this in their robustness checks, noting that:

...to mitigate concerns that results are related to the pandemic, we add a control variable to the regression model to capture the severity of regulations implemented in a player’s home country during the tournament times... Although the aggregate online dummy reduces in size and significance (p-value of 0.172), presumably because lockdowns occurred only during the online tournaments, the effect pattern on the separate tournament dummies remains almost identical relative to the main results...

That doesn't quite allay my concerns, for two reasons. First, it assumes that all players react similarly to the local pandemic context, since it assumes all experience the average effect on their performance. That average effect is statistically insignificant. Second, including the pandemic variable renders the impact of online play statistically insignificant. Part of the problem is that the pandemic is happening at the same time as the switch to online play (for obvious reasons). Clearly, the natural experiment is not sufficient to disentangle the effects of online play from the effects of the pandemic. That really limits what we can learn from this study.

Finally, and interestingly, the negative effects (if we accept that there are some) decrease over time. As Künn et al. note:

...the negative effect of playing online on the quality of moves is strongest for the first (and second) online tournament. Thus, the adverse effect of playing online on the quality of moves decreases over time, possibly because players adapt to the remote online setting...

Perhaps the players have adapted to the online setting, or perhaps they have adapted to the pandemic, or perhaps the pandemic is becoming less severe over time. Given that we can't disentangle the effects of pandemic or online setting, we can't really tell.

I'm not trying to pick on this study, which uses an interesting setting to try and estimate the impact of remote work, in a case where performance can be measured reasonably accurately and consistently. In theory, that should provide as clean a measure of impact as we can find. However, once you recognise the problem in this study, it is easy to see why it would be even more difficult to use the pandemic natural experiment where the data on performance are not as clear.

Saturday, 9 April 2022

Book review: The Race between Education and Technology

In a post last August, I promised a review of The Race between Education and Technology, by Claudia Goldin and Lawrence Katz. After a pandemic-induced delivery delay, and clearing a few other books off my must-read-this-soon list, I've finally finished the book. The easiest way to describe this book is that it is a monograph - essentially, a book-length version of a journal article, with all of the technical detail (and more). It is not really a book written for a general audience. However, while I probably just made it sound negative, that is actually a good thing. Goldin and Katz take the time to fully develop most of their arguments, delving deeply into the data on US education from the 19th Century through until the early 21st Century.

The book's main thrust is an explanation of the changes in inequality in the US over the course of the 20th Century. It is neatly summarised as:

...technological change, education, and inequality... are intricately related in a kind of "race". During the first three-quarters of the twentieth century, the rising supply of educated workers outstripped the increased demand caused by technological advances. Higher real incomes were accompanied by lower inequality. But during the last two decades of the century the reverse was the case and there was sharply rising inequality. Put another way, in the first half of the century, education raced ahead of technology, but later in the century, technology raced ahead of educational gains... The skill bias of technology did not change much across the century, nor did its rate of change. Rather, the sharp rise in inequality was largely due to an educational slowdown.

In the first part of the book, Goldin and Katz review the data on inequality, and then show that skills-biased technological change did not change much over the course of the 20th Century. They then exhaustively review the data on educational change in the US from the 19th Century, through the 'High School movement', and into the later 20th Century where university and college education became the norm for most young people. I learned a lot about the development of the education system in the US, especially the private and public divide in education, both at high school and then at university level. Throughout most of the period, the US maintained a lead in average years of education among its citizenry, compared with other developed countries.

However, as noted in the final chapter of the book, the US has more recently lost its educational advantage, not only in terms of the quantity of education, but also in terms of its quality. Other countries have similar, if not greater, proportion of young people attaining university degrees, and the US is lagging in important measures of educational quality such as the PISA tests of high school reading, mathematics, and science literacy. In contrast with the rest of the book, this section was somewhat underdeveloped. But to be fair, it would require another book to really do the topic justice. Goldin and Katz note that:

Two factors appear to be holding back the educational attainment of many American youth... The first is the lack of college readiness of youth who drop out of high school and of the substantial numbers who obtain a high school diploma but remain academically unprepared for college... The second is the financial access to higher education for those who are college ready.

Although those statements aren't backed by the same depth of analysis and evidentiary support that the rest of the book exhibits, I found them to accord with my own views of the situation in New Zealand as well. Although our education system differs in important ways from the US system, the problems appear to be similar. High schools are not fully preparing students for university education, and there are significant financial barriers that not only stop students from enrolling in university, but also prevent those who do enrol from succeeding to their full potential and maximising their education gains (see my earlier post on this point). The underlying reasons for these problems are not explored in as much detail as they could (and should) be, but Goldin and Katz briefly outline a policy prescription:

The first policy is to create greater access to quality pre-school education for children from disadvantaged families. The second is to rekindle some of the virtues of American education and improve the operation of K-12 schooling so that more kids graduate from high school and are ready for college. The third is to make financial aid sufficiently generous and transparent so that those who are college ready can complete a four-year college degree or gain marketable skills at a community college.

Given the depth of the rest of the book, the policy prescription seems somewhat superficial and underwhelming to me. It would have been nice to have seen how the data supported those proposed policies, or at least a more detailed and robust case made for them.

Nevertheless, despite the final chapter, this is an excellent book, well-written and definitely an exemplar for the comprehensive treatment of historical data, with a strong underlying theoretical model. However, it is worth noting that their more recent update on the research (which I blogged about here), suggests that the theoretical model does not do as good a job of explaining the rise in income inequality in the US in the period from 2000 to 2017. Given that the recent article presents more recent data, for anyone interested in the topic, that article is a better place to start. However, for those wanting to go deeper into the data and the model, this book provides the detail.

Friday, 8 April 2022

Deterring cheating in online assessment requires more than cheap talk

Cheating is a serious problem in online assessments. The move to more online teaching, learning, and assessment has made it all the more apparent that teachers need appropriate strategies and tools to deal with cheating. However, once those tools are in place, they will only deter students if students know that there are cheating detection tools, and if students believe that they will be caught. How do teachers get students to believe they will be caught?

That is essentially the question that is addressed in this new article by Daniel Dench (Georgia Institute of Technology) and Theodore Joyce (City University of New York), published in the Journal of Economic Behavior and Organization (ungated earlier version here). Dench and Joyce run an experiment at a large public university in the US, to see if cheating could be deterred. As they explain:

The setting is a large public university in which undergraduates have to complete a learning module to develop their facility with Microsoft Excel. The software requires that students download a file, build a specific spreadsheet, and upload the file back into the software. The software grades and annotates their errors. Students can correct their mistakes and resubmit the assignment two more times. Students have to complete between 3 to 4 projects over the course of the semester depending on the course. Unbeknownst to the students, the software embeds an identifying code into the spreadsheet. If students use another student’s spreadsheet, but upload it under their name, the software will indicate to the instructor that the spreadsheet has been copied and identify both the lender and user of the plagiarized spreadsheet. Even if a student copies just part of another student’s spreadsheet, the software will flag the spreadsheet as not the student’s own work.

Focusing on four courses (one in finance, one in management, and two in accounting) that required multiple projects, Dench and Joyce randomised students into two groups (A and B):

One week before the first assignment was due, we sent an email to Group A reminding students to submit their own work and that the software could detect any work they copied from another spreadsheet. The email further stated that those caught cheating on the first assignment would be put on a watch list for subsequent assignments. Further violations of academic integrity would involve their course instructor for further disciplinary action. Group B received the same email one week before the second assignment. All students flagged for cheating in either of the two assignments were sent an email informing them that they were currently on a watch list for the rest of the semester’s assignments.

Dench and Joyce then test the effect of information about the software being able to detect cheating, and then they test the effect of students being sanctioned after having cheated on one (or both) of the first two assignments. They found that:

...warning students about the software’s ability to detect cheating has a practically small and statistically insignificant effect on cheating rates. Flagging cheaters, however, and putting them at risk for sanctions lowers cheating by approximately 75 percent.

The results were similar for the finance and management courses, but much smaller for the accounting courses, where the extent of cheating was much lower (interestingly), and where the first projects were due after the sanctions had occurred in the finance and management courses (and so, there may have been spillover effects).

Overall, these results suggest that simply telling students that you can detect cheating and that they will be caught is not effective. Students see it as 'cheap talk' and not credible. Students need to credibly believe that they will be caught, and that there will be consequences. One way to make this credible is to actually catch them cheating and call them out on it. And if this happens in one class, it appears that it can spill over to other classes. And to other semesters, as Dench and Joyce note in an epilogue to the article:

In Finance and Management the rate of cheating after the spring of 2019 was 80 to 90 percent lower than levels reported in the experiment. Given the complete lack of an effect of warnings in the experiment, we suspect that subsequent warnings were viewed as more credible based on the experience of students in the spring of 2019.

Student cheating in assessments is a challenging problem to deal with. If we want to deter students from cheating, we have to catch them, tell them they have been caught, and make them suffer some consequences. Only then will the statements that we make in order to deter students from cheating be credible deterrents.

[HT: Steve Tucker]

Read more:

Thursday, 7 April 2022

What landlords see as important when they set rents

 There is a famous quote, attributed to economics Nobel laureate Ronald Coase, that reads “If you torture the data long enough, it will confess to anything”. Unfortunately, based on my experience this week, that doesn’t appear to be the case. I’ve spent two full days playing with data from a survey a student of mine collected from landlords (members of the NZ Property Investors Federation) back in 2018. The goal was to derive some insights into the factors that landlords see as important when they set rents, and whether those that place a greater importance on tenant attributes are more likely to set rents that are below-market. Unfortunately, I’ve concluded that the data tell us nothing of substance. So, with that in mind, and no prospect of generating a compelling research article from the data, I’ve decided to dump the few interesting bits into this blog post instead.

The genesis of this research was this 2016 post I raised the possibility that landlords offer ‘efficiency rents’:

There are good tenants and bad tenants, and it is difficult for landlords to regulate tenants' behaviour after they have signed the rental agreement. Given this is moral hazard and efficiency wages is one way to deal with moral hazard in labour markets, is there a rental market equivalent of efficiency wages?

First, some context. In ECON100 and ECON110, we discuss moral hazard and agency problems. One such problem is where employees' incentives (after they have signed their employment agreement) are not aligned with those of the employer. The employer wants their employees to work hard, but working hard is costly for the employee so they prefer to shirk. One potential solution to this is efficiency wages (I've previously discussed efficiency wages here). With efficiency wages, employers offer wages that are higher than the equilibrium wage, knowing that this will encourage higher productivity and lower absenteeism from their workers. This is because if workers don't work hard (and avoid absenteeism), they may lose their jobs and have to find a job somewhere else at a much lower rate.

Which brings me to landlords and efficiency rents. As noted above, there is a moral hazard problem for landlords - tenants' incentives (to look after the property) are not aligned with the landlord's incentive (to keep the property in top condition). If the landlord instead offered an efficiency rent (a rent below the equilibrium market rent), then they would have many potential tenants applying for the property, allowing the landlord to pick the best (the least likely to damage the property). It also gives the tenants an incentive to look after the property after signing the tenancy agreement, because if they don't they get evicted and have to find another place to live at a much higher cost.

Maybe landlords offer efficiency rents already and we just don't realise it? 

That’s what we set out to test in 2018. We engaged the NZPIF, and they agreed to support the survey by sending it to their members. We don’t know how big the membership base it (possibly in the thousands), but we had 104 responses to the survey, and 93 of them gave us enough data to be useable for analysis. The median landlord had five properties, and the range was one property to 120 properties.

Do landlords offer below-market rents? Some clearly do (or at least they say that they do). We asked separately about existing tenancies and new tenancies, and 37 out of 93 told us they offer below-market rent to existing tenancies, while 14 out of 93 told us they offer below-market rent to new tenancies. So far, kind of interesting. There were 24 landlords who said that they offered below-market rent to existing tenancies, while also saying that they offered market rent of above-market rent to new tenancies. I took those as indicative of efficiency rents, reasoning that landlords have less imperfect information about existing tenants than new tenants, and so landlords would opt to offer lower rents as they don’t want to lose ‘good’ tenants (this approach has a theoretical basis too – see here).

Unfortunately, it turns out that my measure of efficiency rents is completely unrelated statistically to anything else we know from the survey. Large and small landlords, whether they use property managers, whether they engage in regular rent reviews, the location of the property, etc. are not correlated with my measure. Essentially all I can conclude from that is that whether a landlord offers below-market rent or not is based on unobserved characteristics of the tenant or the property. I guess that is the point of efficiency rents – we don’t observe the quality of the tenant, but the landlord will have discovered some information about tenant quality that we don’t observe. Still, that is pretty unsatisfying as it leaves the survey approach somewhat worthless.

We also asked landlords about what factors (of a total of 21 factors) were important in their rent-setting decisions. We asked these questions in two ways. First, we asked about setting rents ‘on average’ for their properties. We later asked them about a single property, selected a random (the randomisation mechanism here was quite cute – we asked them about the property that is located on a street starting with the letter that is closest to the first letter of their surname [*]), at the last time the property’s rent was set or reviewed. There aren’t systematic differences in the rankings between the two ways we asked (which may again point to idiosyncratic differences in rent setting related to unobserved characteristics of tenants or properties), so I’ll focus on the ‘on average’ results.

We asked the landlords to rate each factor on a five-point scale. Some rated all or most factors important, and others rated all or most factors unimportant, so I standardised the ratings within each landlord, to a measure with a mean of zero and a standard deviation equal to one. Summarising the results for the 93 landlords overall, we get this:

The bars represent how important (on average) each of the factors is. A positive number represents more important on average, and a negative number represents less important on average. The colours of the bars group the factors into different categories (profitability factors; cost factors; local demand factors; property factors; and tenant factors). Overall, on average it appears that the most important factors are the level of rents in surrounding areas, and local demand for rental property (no surprises there). After that, property factors (number of bedrooms, and location and amenities) are important. Least important is local demand from property buyers. That makes sense too. Potential capital gains don’t appear to matter, and property management costs are less important as well (probably because only 43 of 93 landlords used a property manager).

However, the importance of these factors doesn’t appear to differ much based on landlord characteristics (at least, not in a way that makes sense). And the importance ratings are not related to my measure of efficiency rents.

All up, this research didn’t tell us much (at least, not to an extent that makes it publishable other than in a blog post!). That is somewhat disappointing, because there isn’t a large literature on this, and most that exists is theoretical rather than empirical. A better approach for further research might be to look at matched tenant-landlord data, but it’s not clear that such data exists (tenancy bond data is available for New Zealand, but I’m unsure how much data on landlords is captured, or how much data on tenants). I’ll leave that for future work, if I have the energy and inclination (or a motivated student) to work on it again.

*****

[*] This isn’t a perfect means of randomisation, of course. However, I reasoned that approach was better than asking about the property that they last conducted a rent review for (which seems an obvious choice for randomisation). That would be problematic, since the frequency of rent reviews may differ between good and bad tenants, and therefore we would be more likely to receive data on a low-quality tenant or low-quality property. Our approach avoided that problem.