Distrust Your Data

With the launching of 538, Vox and the New York Times’ Upshot, it seems like the age of data journalism is finally here, greeted with both acclaim and concern by media critics. But data journalism is not a new thing. These new sites are just the latest iteration of news applications, which were an iteration of computer-assisted reporting, which was an iteration of precision journalism, all of which are just names for specific techniques and approaches used in the service of reporting the truth and finding the story. In other words, it’s journalism that starts from interrogating the data—and applies the same skepticism and rigor that we apply to the testimony of an expert contacted by traditional phone-assisted reporting.
All of which is to say that data journalism inherits a long tradition of journalists working with data, and that comes with the heavy responsibility to get it right. Specifically, to paraphase something I heard at a NICAR conference once: fear and paranoia are the best friends a data journalist can have. I think about this often when I work with data, because I am terrified about making a dumb mistake. The public has only a limited tolerance for fast-and-loose data journalism and we can’t keep fucking it up.
Critique is always annoying when it’s expressed in indefinite terms. So, I’m going to do something I don’t normally like to do and pick a recent example of a data journalism story gone wrong. This is not to scold those who reported it—indeed, I’m well aware of how easy it is for me to make similar mistakes—but because a specific example provides an explicit illustration of how reporting on data can go wrong and what we can learn from it. And so, let’s begin by talking about porn.
Specifically, a story about online pornography consumption in “red” vs. “blue” states that exploded onto social media a few weeks back. I first noticed it because of a story on Vox that reaggregated an Andrew Sullivan post which in turn reposted a chart made by Christopher Ingraham of the data provided by Pornhub for their study. That chain of links reflects how news spreads online these days, and yet none of those professional eyes caught some glaring flaws in the data.
Before I continue, here’s a brief summary of the findings presented by Pornhub’s data scientists. Pornhub (which is apparently the third most-popular pornography site on the Internet) was approached by Buzzfeed (which is probably the most-popular animated GIF distributor on the Internet) to analyze its traffic and determine whether “blue” states that voted for Obama in the last election consumed more pornography than “red” states that voted for Romney. And so, that’s what the statisticians at Pornhub did, pulling IP addresses from their website’s traffic logs, geocoding their likely locations and deriving a figure of total traffic for each state. They then divided the total hits from each state by that state’s population to derive a hits-per-capita number for each state. As a result, they were able to report that per-capita averages for each state and that blue states averaged slightly more hits per capita than red states.
How To Confuse Yourself With Statistics Unfortunately, the study and the subsequent reporting derived from the Pornhub data serves as a vivid example of six ways to make mistakes with statistics: The first issues begin with the selection of the proxy. In statistics, a proxy is a variable that is used when it’s impossible to measure something directly—for instance, using per-capita GDP as a measure of standard of living. Buzzfeed titled the article about the Pornhub study as “Who Watches More Porn: Republicans Or Democrats?”. Let’s assume that’s the question that Buzzfeed wanted to ask. How would they do it? In an ideal world, they could ask every single Democrat and Republican in the country about their porn watching preferences, but this is obviously unfeasible.


