Overwhelmed by Data? Here’s How to Get Control of It

In this special guest feature, Amnon Drori, Co-founder and CEO of Octopai, discusses how companies can gain visibility and control over their data lineage by leveraging metadata and how this will make GDPR compliance a manageable task instead of an impossible pursuit. Amnon has over 20 years of leadership experience in technology companies. Before co-founding Octopai he led sales efforts at companies like Panaya (Acquired by Infosys), Zend Technologies (Acquired by Rogue Wave Software), ModusNovo and Alvarion, and also served as the Chief Revenue Officer at CoolaData, a big data behavioral analytics platform. Amnon studied Management and Computer Science at the Open University of Tel Aviv.
Big businesses—and smaller ones too—today generate reams of big data. And big data will only get bigger; in 2017, there were nearly 9 billion devices connected to the internet, a number that will exceed 20 billion by the end of this decade, Gartner estimates, and all of them are supplying endless amounts of data.
Indeed, companies have gotten very good at collecting data—but they aren’t as good at using it effectively. It’s estimated that anywhere from 60% to 95% of data collected by companies just lies around collecting dust. But considering that analytics today can do wonders with collected data, providing insights on how to increase sales, what new products to deploy, how to cut administrative or manufacturing costs, and much more, it seems strange that organizations would let the data lie fallow, especially when that data can make a business more profitable. Gaining control of data should be a key business strategy for any organization. So why aren’t businesses in better control of their data?
One reason is that there is just too much data for them to handle. The average GB of data represents roughly 64,782 Word pages, and there are many, many gigabytes to search through. Over 2.5 quintillion bytes of data are produced worldwide each day, and even a mid-sized company today produces far more data than the largest enterprises of the last century. With numbers like those, just structuring the data in databases has become a major problem.
And even when the data is structured, metadata issues can make accessing it a headache. When data is categorized differently in various databases or containers—for example, when some birth dates are categorized European style (year/month/date) and some American style (day/month/year)—searching for that data becomes a major challenge.


