5 Key Components of a Data Sharing Platform

4 min read
Curated from kdnuggets.com →

Read this article for an overview of what the components of a data-sharing platform are

Increasingly, companies are focused on finding ways to connect to new and valuable sources of data in order to enhance their analytical capabilities, enrich their models, or deliver more insight to their business units.   Due to the increased demand for new data sources, companies are also looking at their internal data differently. Organizations that have access to valuable datasets are starting to explore opportunities that will let them share and monetize their data outside the four walls of their business. 

There’s a lot of focus on data sharing. Part of this is because data monetization is becoming a big deal and companies need to figure out a mechanism to get their data in front of people who will buy it. More generally, companies are hungry for data from non-traditional sources that they can use to enhance models, uncover hidden trends, and (yes) find “alpha”. This need has generated a lot of interest in data discovery platforms and marketplaces. Venturebeat featured data sharing specifically as part of their 2021 Machine Learning, AI, and Data Landscape, and highlighted a few key developments in the sphere:

Google launched Analytics Hub in May 2021 as a platform for combining datasets and sharing data, dashboards, and machine learning models. Google also launched Datashare, developed more for financial services and based on Analytics Hub.

Databricks announced Delta Sharing on the same day as Google’s release. Delta Sharing is an open source protocol for secure data sharing across organizations.

In June 2021, Snowflake opened broad access to its data marketplace, coupled with secure data sharing capabilities.

ThinkData, the company I work with, started out in the open data space, and as such has always made data sharing a cornerstone of our technology. Recently, we leaned into data virtualization to help the organizations we work with securely share data (you can read more about that here).

There are some things to consider when it comes to sharing data. The process by which you choose who you will let use your data, and for what purposes, is a strategic question that organizations should think through. But outside of the rules about who will use your data and why, there are also logistical problems involved with sharing data that are more technological than legal in nature. 

When it comes to data – whether it’s coming in or going out, public or private, part of core processes or a small piece of an experiment – how you connect to it is often as important as why. For modern companies looking to get more out of their data, choosing to deploy a data catalog shouldn’t mean you have to move your data. The benefits of a platform that lets you organize, discover, share, and monetize your data can’t come at the cost of migrating your database to the public cloud.

The value of finding and using new data is clear, but the process by which organizations should share and receive data is not. 

Memory sticks aren’t going to cut it, so what does a good sharing solution look like?

Often, when organizations want to share the data they have, they have to slice and dice the datasets to remove sensitive information or create customer-facing copies of the dataset. Aside from the overhead of building a client-friendly version of your datasets, this is bad practice from a data governance perspective.

Making and distributing a copy of a dataset creates a new way for data and data handling to go wrong every time, and opens up possibilities for inconsistencies, errors, and security issues. If you want to share the dataset to ten different people for ten different use cases, you’re repeating work and multiplying risk every time you do.

A good data sharing solution lets data owners create row–and column-level permissions on datasets, creating customized views of one master file, ensuring all data shares still roll up to a single source of truth.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at kdnuggets.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.