The 3 Vs and Unstructured Data Analytics

3 min read
Curated from tdwi.org →

Becoming more efficient in all aspects of unstructured data management is key to your analytics programs.

Enterprise data volumes are blowing through the rafters, consuming a significant portion of the IT budget. More than half of IT leaders report that their organizations are managing 5PB or more of data and most (68 percent) are spending more than 30 percent of their IT budget on data storage, backups and disaster recovery, according to the Komprise 2022 State of Unstructured Data Management, a third-party survey conducted earlier this summer and sponsored by my company.

Five petabytes is a lot of data (about 1.25 billion digital photos’ worth, for example) and much of it is unstructured, meaning that it doesn’t fit neatly into rows and columns in a database. This unstructured data — such as log files, IoT sensor data, microscopic data, application data, user documents, manufacturing test data, and medical images — is an untapped gold mine for primary research and analysis.

Traditionally, data analytics has relied on data warehouses and mining structured and semi-structured data from spreadsheets or financial documents. Yet this data is just the tip of the iceberg, considering that an estimated 80 percent or more of global data is unstructured. With advances in cloud computing, machine learning (ML), and AI tools, unstructured data analytics is now a prime opportunity. ML algorithms depend upon large quantities of data and the cloud delivers a wide array of on-demand, compute-intensive services to affordably run big data infrastructure and analysis like never before.

Today there are a multitude of cloud-based ML and AI services for different use cases — from image and audio pattern recognition to personally identifiable information (PII) identification. Some interesting and valuable use cases for unstructured data analytics include medical insurance fraud detection, autonomous vehicle testing, malicious actor detection, precision medicine, and customer sentiment analysis of call center audio files.

The Komprise survey showed that 65 percent of organizations plan to or are already investing in delivering unstructured data to their new analytics/big data platforms. To be successful in unstructured data analytics, you must jump through several hurdles compared to the relatively straightforward process of mining structured data in databases and spreadsheets. Gartner analyst Doug Laney introduced the 3Vs concept in a 2001 MetaGroup research publication, 3D Data Management: Controlling Data Volume, Variety, and Velocity. When it comes to unstructured data, these challenges include:

Volume of data. Because there is so much data in organizations today, you can’t feasibly analyze it or copy it all to a cloud service or big data platform.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at tdwi.org →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.