Skip to content
7wData Data and AI tools, companies, events, podcast
  • Tools
  • Companies
  • Podcast
  • Articles
  • Events
  • Newsletter
  • Research
  • Sponsor

Table of Contents

Artificial Intelligence 2017 • By Yves Mulkers

MapR takes a stab at data governance, in an age of data anarchy

MapR takes a stab at data governance
3 min read
apache software foundation, API, Big Data
Curated from zdnet.com →

Data governance — the discipline of inventorying and annotating your data sets, determining their accuracy, pedigree and quality and properly securing them — is an important focus area for the industry. In the conventional database world, Enterprise Information Management (including ETL, data quality management and master data management) has addressed these needs for some time. In the data lake world, though, efforts have been far less earnest.

Granted, there are data catalog products, and lineage products. There are various security/access control solutions and there are metadata management systems as well. Cloudera has its Navigator product, and there’s even open source Apache Atlas (incubating), borne of a Hortonworks project called the “Data Governance Initiative.” Some analytics products even have governance features of their own, enticing customers away from bringing in another vendor and platform to handle governance requirements.

On Tuesday, MapR announced a data governance initiative of its own. It is comprised of an interesting architectural approach, a few key partnerships, and a services offering to go with it all. I’ll first detail MapR‘s announced offering, and I’ll conclude with some analysis (aka a rant) on the state of data governance in the Big Data world.

Topic: PreprocessorOn the technical side, MapR has come up an approach that, to me at least, is pretty novel and clever. The company has taken a prescriptive stance here and is advising customers that all data ingestion should go through MapR Event Streams (MapR-ES, formally known just as MapR Streams), the company‘s Kafka API-based publish/subscribe platform for handling event-based data ingest.

The hook, as it were, is this: by configuring a pre-processor on the MapR-ES topic, all data pushed through it can be observed, its discernible metadata captured in a MapR-DB document database, and metadata changes can also recorded there. This allows for metadata cataloging and, if derivative data set creation is managed similarly, and all MapR-ES events are retained, data lineage can be determined comprehensively, just by “playing back” the events.

The partner part So MapR provides the raw infrastructure to get metadata and lineage information. But it doesn’t offer a data catalog facility that would let data lake users search for data sets, tag them, see which of them are certified and see star ratings for them, provided by other users.

That’s where partners and their products come in. Waterline Data and Collibra, each of which offers data catalog and data lineage functionality, are key partners. Cask, whose Data Application Platform (CDAP) provides a unified API over various Big Data components, and specific APIs for metadata inspection and for audit, is a partner as well.

By themselves, each of these products only catalogs what’s entered into them. They work as long as everyone uses them (or codes to them, in the case of CDAP).

In the 7wData directory

Get the AI & data signal, daily.

48k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

Compare the tools & companies behind this topic

Browse the directory →
  • ChatGPTTool
  • AgentforceTool
  • Scikit-learnCompany
  • OONNX RuntimeTool
  • PyTorchCompany
  • TensorFlowCompany
  • PycaretCompany
  • MMxnetCompany

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at zdnet.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.

Want the structural read on any AI or data company?
INS7GHTS

Want a sharper read on this topic?

Ask ins7ghts how the players compare, what people are actually shipping with, and where the trade-offs land.

Tweet LinkedIn Bluesky Threads Email

Related Articles

What Changes When AI Is So Accessible That Everyone Can Use It?
Artificial Intelligence

What Changes When AI Is So Accessible That Everyone Can Use It?

3 min read • 2018
How to calculate cloud migration costs before you move
Artificial Intelligence

How to calculate cloud migration costs before you move

4 min read • 2021
Blockchains aren’t just tech
Data Management

Blockchains aren’t just tech, they’re new economic systems

3 min read • 2018
7wData

Independent reporting on AI and data: daily newsletter, podcast, deep dives.

Read

  • Ins7ghts newsletter
  • AI Beat newsletter
  • Latest articles
  • Podcast
  • Research guides

Use

  • Tools directory
  • Company directory
  • Research
  • Events
  • ins7ghts

Company

  • About
  • Contact
  • Sponsor a slot
  • Media kit
  • RSS feed

Follow

  • LinkedIn
  • X
  • YouTube
  • Instagram

© 2026 7wData. Independent. Belgium-based.

Privacy Cookies Terms Imprint Cookie settings
INS7GHTS
Cookies on 7wData

We use strictly necessary cookies for the site to work, and optional analytics cookies to understand how readers use 7wData. We never share your data with advertisers. See our Cookie Policy.

Get the AI & data signal, daily. 48k+ already do.
Thanks. Check your inbox to confirm.
Get the AI & data signal

One curated email a day. 48k+ data & AI professionals already read it.

No spam. Unsubscribe anytime.

Check your inbox.

We just sent a confirmation. Click the link to start receiving the daily signal.