
Designing, Operating and Managing an Enterprise Data Lake
About This Event
This seminar explores a new approach in designing, building and managing an enterprise data lake to get control of your data. This includes introducing a data refinery and information catalog to produce and publish enterprise data services for consumption across your company.
The Enterprise Data Lake seminar
This 2-day seminar looks at the challenges faced by companies trying to deal with an exploding number of data sources, collecting data in multiple data stores (cloud and on-premises), multiple analytical systems and at the requirements to be able to define, govern, manage and share trusted high quality information in a distributed and hybrid computing environment. It also explores a new approach of how IT data architects, business users and IT developers can collaborate together in building and managing an enterprise data lake to get control of your data. This includes data ingestion, data discovery, data profiling and tagging and publishing data in an information catalog. It also involves refining raw data to produce enterprise data services that can be published in a catalog available for consumption across your company. We also introduce multiple data lake configurations including a centralised data lake and a ‘logical’ distributed data lake as well as execution and governance across multiple data stores. It emphasises the need for a common collaborative process and common approach to governing and managing data of all types.
Learning objectives
- How to define a strategy for producing trusted data as-a-service in a distributed environment of multiple data stores and data sources
- How to organise data in a centralised or distributed data environment to overcome complexity and chaos
- How to design, build, manage and operate a distributed or centralised data lake within their organisation
- The critical importance of an information catalog for delivering data-as-a-service
- How data standardisation and business glossaries can help define the data to make sure it is understood
- An operating model for effective distributed information governance
- What technologies they need and implementation methodologies to get their data under control.
- How to apply methodologies to get master and reference data, big data, data warehouse data and unstructured data under control irrespective of whether it be on-premises or in the cloud.
