Data Lakes Still Need Governance Life Vests

2 min read

As a central repository and processing engine, data lakes hold great promise for raising return on data assets (RDA).  Bringing analytics directly to different data in its native formats can accelerate time-to-value by providing data scientists and business users with increased flexibility and efficiency.

But to realize higher RDA, data lakes still need governance life vests.  Without data governance and integration, analytics projects risk drowning in unmanageable data that lacks proper definitions or security provisions.

Success with a data lake starts with data governance.  The purpose of data governance is to ensure that information accessed by users is consistently valid and accurate to improve performance and reduce risk exposure.

Data governance is a team sport.  Collaboration among and between data scientists, IT and business teams define the use cases that the architecture and analytics software will support.

Understanding business and technical requirements to identify the value data provides is the first step to developing a data governance cycle.  Data governance establishes guidelines for consistent data definitions across the enterprise.  It also defines who has access to specific data and the purposes of usage.

Without data governance, it’s impossible to know whether the information presented is accurate, how and by whom it has been manipulated.  And if so, with what method, and whether it can be audited validated or replicated.  As departments maintain their own data – often in spreadsheets – and increasingly rely on outside data sources, a verifiable audit trail is compromised, exposing the firm to compliance violations.

Including security teams in data governance is also crucial. By understanding what data will be brought into the data lake and the user access permissions, security teams can better understand potential risks.  They can build stronger protection around critical data assets while becoming more resilient and responsive to incidents.

As the data lake becomes the repository for more internal and external data, IT must integrate the data lake with the existing infrastructure.  One of the benefits of a data lake is that it can ingest data without a rigid schema or manipulation.  Integration reduces errors and misunderstandings, resulting in better data management.

Modern data integration technologies automate much of the data quality, cataloguing, indexing and error handling processes that often encumber IT teams.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at bigdatanews.datasciencecentral.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.