WAIT!! WHAT?!?!: āMr Albouy Reaches His Conclusion by Omitting Half the Data from the Original Sampleā: ANOTHER CHART OF THE DAY
The Economist published an unbylined, unsourced piece asserting Daron Acemoglu ācounters that [David] Albouy reaches his conclusion by omitting half the data from the original sample.ā That claim is simply wrong, and Albouy is right to be angry ā no fact-check, no source, no context. Albouyās actual point is narrow and correct: AJR have no real first stage. Once you correct for clustering, drop 36 conjectured mortality rates, and control for barracks-versus-campaign sources, the mortalityāexpropriation relationship collapses toward one-in-three significance. The second-stage test statistic isnāt a t-distribution; itās near-Cauchy ā infinite variance, no mean. The IV estimates are unreliable. Moreover, Acemoglu, Johnson, and Robinson ought to be very grateful to David If you take their IV results seriously, the effects implied are embarrassingly and implausibly large: a clock that chimes thirteen. Albouy provides an explanation for what is otherwise a very large implausibility in their story:
One cannot know what to make of this paragraph in the London Economist:
Anonymous: The Worldās Most Influential Economist Is Oddly Unconvincing : āDavid Albouy⦠showed that some countries were assigned mortality rates borrowed from other[s]ā¦. Correct⦠and the [Acemoglu] paperās estimates become unreliableā¦. Buchner⦠and colleagues reported that experts they surveyed were somewhat more likely to side with Mr Albouy. Mr Acemoglu⦠counters that Mr Albouy reaches his conclusion by omitting half the data from the original sample, including on important countries like America, Canada and Australia. It is this combination, along with some statistical choices, that introduces the unreliability, he says. He adds that if he were redoing the paper today, he would make āa number of changes, including in some of the estimation detailsāāthough not to the mortality dataā¦
To start with, the Economistās lack of bylines makes hit pieces like this one on Daron Acemoglu unconvincing. The lack of sources does as well. Normally, you expect an unsourced āsaidā or ācountersā to be something said to the reporter. But the story quotes a tweet from Noah Smith:
Iāve been yelling about Acemoglu for literally a decadeā¦
And it did not contact Noah. It quotes a podcast segment from Larry Summers:
He leaves out entirely in that analysis the possibility that we will have more rapid scientific progress, more rapid social-scientific progress, or better decision-making because of artificial intelligenceā¦
without stating the source as well.
Thus I have no idea what the context of the part of the story I take exception toāthe statement that Acemogu ācounters that Mr Albouy reaches his conclusion by omitting half the data,,, along with some statistical choicesā¦āācomes from. Is this something that Daron said to the story-writer? If so, this is a very bad thing to say. Is the context otherwise? I would like to see the context.
In any event, that claim is simply wrong. And David Albouy is right to be seriously pissed off:
David Albouy: āThe funny thing is that the article is about how Daron is isnāt really trusted. And then he ended up making a lie about my work in the article itself. And the reporter didnāt bother fact checking it or running it by me eitherā¦
Here is David:
The question is what to make of these two charts:
So let me turn the microphone over to David:
David Albouy: The Colonial Origins of Comparative Development: An Empirical Investigation: Comment : āThere are several reasons to doubt the reliability and comparability of their European settler mortality ratesā¦.
First, only 28 countries have mortality rates that originate from within their own borders. The other 36⦠are assigned rates based on conjectures the authors make as to which countries have similar disease environments. These assignments are generally unfounded and potentially contradictory.⦠At a minimum, the sharing of mortality rates across countries requires that statistics be corrected for clustering (Moulton 1990). This correction alone noticeably reduces the significance of the results. If, in the hope of reducing measurement error, the 36 conjectured mortality rates are dropped from the sample, the point estimates relating mortality rates with expropriation risk become substantially smaller, particularly in the presence of covariates, which often gain significance.
Second, the mortality rates never come from actual European settlersā¦. Instead, the data come primarily from European and American soldiers in the nineteenth century⦠[some] at peace in barracks⦠others⦠on campaignā¦. Controlling for the source of the mortality rates weakens the empirical relationship between expropriation risk and mortality rates substantially. Furthermore, if these controls are added and the conjectured data are removed, the relationship virtually disappears, suggesting that it is largely an artifact of the dataās constructionā¦.
Without a robust relationship between expropriation risk and mortality rates, the AJR IV estimates of the effect of expropriation risk on GDP per capita suffer from weak instrument problems: point estimates are unstable, and corrected confidence intervals are often infiniteā¦.
Albouyās point is this: AJR do not have a first-stage for their instrumental-variables regression. As I put it, when I teach the paper, they report that the association between log mortality, as they assign it, and perceived expropriation risk in the 1980s ( Never mind what extent perceived expropriation risk by a consulting firm evaluating political risk can be understood to be a measure of institutions) passes the null-hypothesis test at 1/1000, at one chance in a thousand. They ought toābecause their mortality assignments are carrying information about continents, which carry a lot of information that makes expropriation risk look more salient than it is in the IVāhave reported column (4), which passes the null-hypothesis test at 1/25. And by the time one has noted that some soldiers are in barracks and some of the data come from poor laborers, the first-state is down to 1/3. Thus the distribution of their test statistics from the second-stage of their IV regression is not a t-distribution, but is because of the weak-instrument problem near to a Cauchy distribution: that thing that not only has infinite variance and standard deviation, but does not even have a mean.
That is a fair point. The IV estimates are highly unreliable.
After pointing that out, I go back to the OLS correlation between expropriation risk and prosperity today, and we talk about the various ways the world might work that would produce that OLS scatterplot:
What other than a relationship between trust in the rule of law and the solidity of property rights on the one hand and economic activity leading to prosperity on the other hand could produce this scatterplot?
Now Acemoglu, Johnson, and Robinson defend the hill of their IV because, according to the rules of economics since the empirical-causal turn, an OLS scatter cannot be interesting or the basis for an AER paper. But their OLS scatter is interesting, and important, and worth talking about.
Daron Acemoglu, Simon Johnson, and James Robinson: Hither Thou Shalt Come, But No Further: Reply to āThe Colonial Origins of Comparative Development: An Empirical Investigation: Commentā : āOverall, Albouyās āCommentā amounts to a series of objections to our approach. All of these objections, upon closer inspection, are far from compelling, are often unfounded, and prove minor and largely inconsequential for the robustness of our results. The big picture from AJR (2001) remains intact and remarkably robust: Europeans were more likely to move to places that were relatively healthy, and when they moved in larger numbers, they imposed better institutions, which have tended to persist from the colonial period to today.
In my view, this rhetorical position by Acemoglu, Johnson and Robinson is a huge mistake. They should be very glad to drop their IV results from the discussion.
Acemoglu, Johnson and Robinson went down the road of their IV because of the potential criticism that causation is not flowing from perceived security of property to prosperity, but rather that prosperity has lots of effects that lead to perceived security of property. They thought that they could use settler mortality to identify a component of perceived expropriation risk today that was plausibly independent of these reverse causation factors. In my view, the right way to understand what they wound up doing is to say that they regressed current prosperity on what perceived expropriation risk today would be if you knew only about settler mortality in the past and nothing else. The slope of that regressionāthe regression of prosperity today on what one would expect perceived expropriation risk to be from settler mortalityāis the hill they are willing to die on.
Their OLS scatterplot gives them a coefficient β-hat =0.522: a one-unit increase in their security-of-property today index leads to an 0.522-unit increase in prosperity today, a 68.5% increase.
And their IV scatterplot gives them a coefficient βā-hat = 0.944, which means a one-unit increase in what you would think perceived expropriation risk is today based on settler mortality leads to an 157% increase in prosperity today.
In the real world, a country like New Zealand with a log prosperity score of 10 and a perceived security-of-property score of 10 was back before 2000 about 5 times richer than an Egypt with a log prosperity score of 7 and a perceived security-of-property score of 7. AJRās IV results say that the effect of security-of-property on prosperity ought to be much much bigger than that: it ought to be about 20 times richer. And it would be, if the relationship between security of property leading to investments in physical and human capital and in enterprise and innovation were the only things operating, were there no reverse causation by which prosperity leads to social unrest that threatens the security of property. But there is such social unrest, AJRās IV results tell us. There is powerful reverse causation: a doubling of prosperity sets in motion societal forces that would, if they were the only things operating, reduce your perceived security-of-property score by 0.4.
Now: any claim that prosperity structurally reduces governance quality contradicts an overwhelming body of evidence across multiple disciplines:
1. The Modernization Hypothesis (Political Science): Lipset (1959) argued the that it was economic development that created the social conditions for democracy and good governanceāan educated middle class, urbanization, organizational capacity, and norms of civic participation. Przeworski and Limongi (1997) and Boix and Stokes (2003) found strong empirical support that higher income increases the probability of democratic consolidation. Acemoglu et al. themselves, in later work (Journal of Political Economy, 2008), examine this relationship. The literature is not unanimous, but no serious strand of it argues that prosperity destroys governance quality.
2. Historical Evidence on Expropriation and Revolution: Revolutions, expropriations, and institutional breakdown have historically occurred overwhelmingly in poor countries, not rich ones. The Russian Revolution and the wave of postcolonial nationalizations all arose in contexts of poverty, grievance, and institutional fragility, not in contexts of prosperity. The handful of rich-country institutional collapses (Weimar Germany) are exceptions, and even there the mechanism was economic collapse, not economic success.
3. The Resource Curse Literature: The closest empirical phenomenon to is the resource curse: oil wealth sometimes corrodes governance by enabling authoritarian consolidation without taxation. But this mechanism is specific to extractive resource wealth, not to prosperity in general, and it operates through a very particular political economy channel (Ross, 2001; Robinson, Torvik & Verdier, 2006). It cannot be generalized to a structural negative relationship between income and governance across all former colonies.
4. The Cross-Sectional Evidence: Simply looking at the data: the richest former coloniesāSingapore, South Korea, Botswana, Mauritiusātend to have better governance, not worse. The poorestāthe DRC, Haiti, South Sudanātend to have the worst. A relationship is not what the scatter plot shows.
Thus the implied backwards causation from prosperity to poor governance quality is not merely theoretically awkwardāit is empirically falsified by the entire body of comparative political economy.
This means the only semi-live explanation of the OLS-vs-IV gap isāif we take it as anything other than the Cauchy distribution arriving at the picnic and getting out its refreshments, as an example of play, stupid games and win stupid prizesāthat it is the result of measurement error in governance quality attenuating the OLS estimate and pushing it downward. That is the interpretation AJR themselves favor. They conclude that the instrumental-variables strategy does not primarily correct for reverse causation. Rather, it primarily corrects for attenuation bias from mismeasured institutions.
But taking that governance quality is mismeasured gets Acemoglu, Johnson and Robinson into even worse trouble.
Their argument is that we cannot today see the true institutional quality of governance G. Instead we see a corrupted and noisy measure of institutional quality H. But if we look at just that part of H that is correlated with settler mortality, we recover a better measure of the institutions that matter. In order for this explanation to work, however, settler mortality in the age of imperialism has to (a) be correlated with that part of modern-day institutions that matter for prosperity, the G, while also (b) being uncorrelated with those parts of modern-day institutions that do not matter for prosperity. Settler mortality C must be: 1. relevant, in that it is correlated with the component of observed institutions that genuinely cause prosperity; and 2. selectively orthogonal, uncorrelated with the part of observed institutions counted as good governance that do not in fact matter for prosperity.
And my response to this is simply this: You are putting me on.
There is no way that 1800s soldier mortality gives us a better measure of institutional quality today, than does looking around at everything we can see about institutions and constructing high-information measures.
The right response from Asimov, Johnson, and Robinson to Albouy is not to dig in deeper, is not to die on the hill of a weak instrument IV, but rather to be very grateful that Albouy has pointed out the magnitude of the weak instrument problem and thus provided a Cauchy distribution explanation of the weirdly implausibly and embarrassingly large size of their IV coefficient. They really do not want to be on either of the hills that defending their IV could put them on:
either prosperity has catastrophically destructive effects on governance,
or (ii) colonial-era mortality is a better measure of what matters for prosperity today in modern institutions than modern institutional analyses can produce.
Neither of those is at all defensible.
But Wait! There Is MOAR! Much, Much Moar!!
ADDITIONAL CRITIQUE 1: Itās Human Capital, Not Institutions: Journal of Economic Growth, Glaeser, La Porta, Lopez-de-Silanes, Shleifer 2004: Europeans who settled in healthy places brought themselvesāeducated, literate people with specific cultural and legal traditions. The AJR IV identifies where Europeans settled, but settler presence = human capital accumulation, not just institution-building. The instrument may be picking up human capital rather than institutional quality.
But: AJR control for fraction of European population directly, and the institutional effects survive. Moreover, human capital and institutions are not cleanly separable ā colonists built institutions deliberately. The critique identifies a real confound but doesnāt break the core result.
ADDITIONAL CRITIQUE 2: Geography, Not Institutions: Various, Sachs & al.: Malaria, latitude, and disease burden directly affect economic productivity, not just through the institutional channel. Settler mortality may proxy for the disease environment that directly depresses output today, violating the exclusion restriction.
But: AJR control for latitude, temperature, humidity, malaria prevalence, and moreāand the institutional effect survives all of these. The paperās robustness tables are unusually thorough on exactly this point.
ADDITIONAL CRITIQUE 3: Institutions Donāt Persist That Way: Various, Path Dependence Skeptic Historians & Political Scientists: The persistence mechanism is assumed rather than demonstrated. Colonial institutions were often radically transformed at independence, and many countries have undergone multiple regime changes. Why would 17th-century institutional choices echo so cleanly into 1990s PRS scores?
But: The correlation between early institutions and current institutions is actually quite well-documented empirically. The persistence is real, even if the mechanisms are complex. AJRās own later work, i.e., Why Nations Fail, develops the persistence story more carefully.
ADDITIONAL CRITIQUE 4: The Exclusion Restriction Is Untestable: Standard IV Critique: Settler mortality back then affects current income through channels other than āinstitutionsāāby shaping culture, work norms, trust levels, or through direct epidemiological legacies.
But: Untestability is a feature of all IV designs, not a special problem for AJR. And AJRās robustness to controlling for an unusually long list of potential direct channels is about as good as IV evidence ever gets.
Bottom line (for teaching): The Albouy data critique is the one that genuinely stingsānot because it refutes the finding, but because it raises legitimate questions about fragility. On the other hand, the very large size of the estimated IV coefficient is very embarrassing and suggests that something has gone catastrophically wrong with the analysis. Albouy provides a way out of that very legitimate critique.
The others are important conceptual challenges that AJR largely anticipated and addressed. The paperās core claimāthat colonial history explains a large fraction of income differences today with āinstitutionsā as a primary channelāhas proven surprisingly durable; while also remaining puzzling and, in a sense, unbelievable.
How it works
Once you click Generate, Ollama reads this article and crafts 5 comprehension questions. Your answers are graded against the article content ā general knowledge won't be enough. Score 70+ to count toward your certificate.
Questions are cached ā you'll always get the same 5 for this article.