The conflict between data science and cybersecurity
Big data analytics, machine learning and predictive analytics is supposed to be the panacea that will allow businesses and organizations to solve to all their problems. The promise is that given access to large swath of data, individuals in an organization can use new and interesting ways to solve complex problems.
When it comes to big data and data mining, the more data you have the more accurate your analytics.
The problem with this is that it assumes everyone can (and more importantly should) have access to all the information in an organization. In fact, security and privacy concerns means that the exact opposite is true.
Cybersecurity is in direct conflict with the basis tenants of data analysis, especially big data analysis, predictive analytics and data mining. Data analysis, especially big data, is about opening access to a lot of the data in an organization to find new some interesting ways to solve business problems. For example, giving a marketing team information that they typically don’t use so they can find new and interesting ways of positioning your products or do some predictive analytics for sales forecasting.
Data Governance and security is about locking down access to data so that only the individuals who should have access get it. It’s about granting the least privileged access and pruning open access to data. This is becoming all too important these days as we see massive breaches where customer or employee data is lost and companies face, at minimum, reputational losses. At worse, they see massive fines, or loss of profit.
It’s not just for data breaches, regulatory requirements also play a big role is who should have access to what data. Across the EU we are seeing regulations that allow customers to be forgotten. This means you must have a complete understanding of where your sensitive data lives and have the ability to track down and delete all instances of a single customer’s information.


