The hidden value in imperfect big data

3 min read
Curated from techrepublic.com →

For companies that build their competitive advantage on big data analytics, data sources are everything. If you are putting garbage into your high-powered analytic system, there’s not much value that can come out.

When defining big data for the purposes of a competitive strategy, I contend that big data should be freely available with no obligations or constraints. However, the price you pay with most freely available data is uncertainty. Great big data strategists do, in fact, consider their sources very carefully and learn how to live with this uncertainty.

The axiom that you must have pristine data sources was challenged once the era of big data analytics emerged. We’ve had the idea of fuzzy math in the artificial intelligence world for quite some time, but it never applied to the business intelligence and data warehousing world until the data science movement. Now, fuzzy data sources aren’t a categorical reject—they’re something that deserves careful consideration.

For instance, with the aid of recent big data analytic technology, many companies are combing through social media feeds to conduct sentiment analysis. There is no possible way to get a high level of data quality from a Twitter feed; however, it’s still very useful, given you’re clear on its confidence rating—a very important concept in profiling data sources. A confidence rating is an overall assessment (typically in percentage terms) of your data source’s quality.

Without the concept of a confidence rating, you’re likely to over-cleanse your data. It’s easy to arbitrarily consider imperfect data sets or sources as invalid, and exclude them from your analytic system. Social media feeds aren’t the only place where you can make this mistake—unstructured data that lies in documents, video, and audio is all fair play now. Instead of ignoring it, map the data through with a confidence rating that provides a disclaimer for what you’re analyzing.

Your transformation maps with fuzzy data sources must be tight, though. I approach data mapping in these situations like I approach audit-proofing. Imagine your system could be audited at any given time, and you must be prepared to substantiate and explain all transformations from your sources to your targets. To do this, there must be a logical trace from your target data back to your—sometimes fuzzy—source.

Expert systems are the best way to handle this situation. Even if you have a transformation tool like Informatica to explain the logic of the transformations, in most cases it won’t explain the reasoning behind the logic.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at techrepublic.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.