You won’t clean all that data, so let AI clean it for you

Don’t cheat and Google the answer to this question:
How many zeros are in a zettabyte?
A zettabyte is one sextillion bytes and one of them is a tenth of all data created each year. Picture the number 10 followed by 21 zeros. If that sounds like an unnecessary factoid, you would be forgiven.
The reason you should care is that market intelligence firm IDC says that by 2025 we will create 180 zettabytes of data annually. This massive acceleration of data creation means that all the enterprise efforts to make use of their data will continue to grow exponentially. It is time to give up on trying to organize and clean it all.
A data architect friend says that every enterprise’s data is great until you try to use it. Analytics, business intelligence, systems integration and collaboration projects tend to be envisioned and approved based on great business cases. But they also reveal the scary realities of enterprise data.
So, it is common for there to be a data cleaning project before the project that the business really wants done. These projects are often brute force, trial-and-error approaches to cleaning data and the level of uncertainty, lack of big data skills and the limitations of the tools meant these efforts tend to be high in expense and headache but often light in measurable business ROI.
Unlike other data-centric initiatives, instead of pitching a project to corral and clean your data before launching your artificial intelligence initiatives, instead use machine learning to get you there faster and easier.
Traditional enterprise data strategies suffer from a central flaw of scalability. The more data you throw at them, the more that the inconsistencies in the original data collection, storage and manipulation come through. So, you must keep developing new rules for how to handle this. It’s the worst kind of project scope-creep because, too often, you realize that every time you find a new exception, you are getting farther from your goal, not closer to it. Essentially, you have trapped yourself into trying to spend money to clean data faster than your enterprise can generate data.
Most data will always be unstructured and the amount of user-generated data, or raw material, is increasing constantly. It’s amazing what has been accomplished in the past just by using SQL.
I have worked on some massive, internet-scale data platforms which included user-generated content from sixty million people a month, feeds of data from sensor devices, a search engine and social media platforms.


