What are Seven Types of Big Data Debt

2 min read
Curated from blogs.starcio.com →

As more organizations have embarked on agile software development over the last five years as part of digital transformation programs, the term technical debt is more widely understood. Teams that develop code leave artifacts behind that require improvements, reengineering, refactoring, or wholesale rewriting.

Some technical debt is done purposefully to deliver applications faster, while other forms of technical debt emerge over time with age and increased usage. Developers focus on fixing technical debt in their code. At the same time, CIOs often apply the term at a macro level and include legacy systems, monolithic applications, or any technology that needs an architecture upgrade.

But what about data?
More organizations are trying to be data-driven , invest in self-service BI programs , or want competitive advantages using machine learning and AI . How much data debt inhibits accurate, reliable, and unbiased analytics?

How Should Organizations Define Big Data Debt?

Let’s try to put some levels of big data debt:
Dark data is data that hasn’t been cataloged or thoroughly analyzed . It is the lowest form of big data debt because it represents data mostly unknown and unused in analytics or machine learning experiments.
Dirty data is data that has known and often unknown data quality issues. This includes fields that need to be normalized or are missing values across a large percentage of data. Data quality and profiling tools are potential options for cleansing dirty data.

Duplicate data comes from data sources that are full or partial copies of primary data sources. They represent everything from spreadsheets to database temporary tables that were created for single or temporary use and often have derivative data added to them. 

Murky data is complex data that is only used by a small number of business analysts, data scientists, or citizen data scientists in the organizations. Murky = not well documented or understood and so subject matter experts are required to interpret the data before it’s applied accurately. Building data catalogs and defining data dictionaries are two ways to address murky data. 

Dysfunctional data occurs when the current data management tools are the wrong ones or are poorly structured to enable optimal use of the data in applications, data analytics, or machine learning. Data stored in file systems, unstructured data stored as CLOB database fields, media data that hasn’t been tagged are all examples.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at blogs.starcio.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.