How to work with someone else’s data

4 min read
Curated from medium.com →

If you’re about to jump on the citizen data scientist bandwagon (diving into COVID-19 data, perhaps?) then there are a few things you should know about dataprovenance…

Society is plagued by distorted expectations regarding data, littered with nonsense like “numbers can’t lie” and “it’s just your opinion until you show me the data” (no, it’s still your opinion) and “I looked at data, so now I’m informed.”

There comes a time in every child’s life when they must learn that: 1) The tooth fairy isn’t real. 2) Things don’t just magically work out because you have some numbers. It really matters where those numbers came from. (Some children are decades overdue for this developmental milestone.)

Anyone can put some electronic scribbles in a table and call it data. That doesn’t make it good/true/useful/worthy in the sense you associate with Science.

Even if the dataset was collected carefully, are you sure you know what happened to it on its way to you? The only reason that a villain won’t remove inconvenient rows from your dream dataset (“Hide data from you? I would never! Those were outliers.”) or aggregate things in a way that skews the message before sharing data with you is that it’s too easy. (Self-respecting villains prefer a challenge.)

Ability to reason about provenance is one of the basic requirements of data literacy, so let’s talk about a few distinctions:

and then round out the article with advice on how to work with inherited data.

You’re using primary data if you (or the team you’re part of) collected observations directly from the real world. In other words, you had direct control over how those measurements were recorded and stored.

What’s the opposite? Inherited (secondary) data are those you obtain from someone else.

Inherited (a.k.a. secondary) datasets are like inherited toothbrushes: an act of desperation. You’d always prefer to use your own if possible. Unfortunately, you might not have that option. Collecting primary data can be prohibitively expensive.

Get the AI & data signal, daily.

335k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

While primary data carries a whiff of superiority reminiscent of artisinal cheese, anyone who insists that you’re worthless for chowing down on inherited data should check their privilege. Individuals (as well as firms without strong data traditions or newcomers to a domain) might not have the resources to collect data on their own, especially when the project requires very large datasets or specialist skills/equipment. For example, not everyone can afford to take pathogen measurements in a lab environment with a high biosafety rating.

Buyer beware: If you’re forced to work with someone else’s data, don’t be surprised later on if things don’t pan out the way you’d hoped. There’s no guarantee that inherited data will serve your needs… (Also, the tooth fairy isn’t real.)

Here’s a quote I like from “the father of statistics”, R.A. Fisher:

That goes for your inherited dataset too — its collection was finished before you arrived on the scene. As far as your intended purpose goes, it might be kaput and the most you can squeeze out of it is an autopsy. I hope you’ve taken enough journeys around the sun not to let this ruffle your feathers.

If you’re on the verge of protesting that it’s better to have some data rather than no data at all, replace the word “data” with “noise” / “lies” / “distractions” and try your sentence again. Quality is everything, and if you didn’t collect the dataset yourself, you have no control over what was measured and (perhaps more importantly) what was left out.

When you’re forced to work with inherited data, you’ve got four main problems to worry about:

Time for some additional jargon along these lines.

Captured data are intentionally created for a specific analytical purpose, while exhaust data are byproducts of digital/online activity. Exhaust data usually come about when websites store activity logs for purposes — such as debugging or data hoarding — other than specific analyses.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at medium.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.