The essential check list for effective data democratization
Truly data-driven companies see significantly better business outcomes than those that aren’t. According to a recent IDC whitepaper, leaders saw on average two and a half times better results than other organizations in many business metrics. In particular, companies that were leaders at using data and analytics had three times higher improvement in revenues, were nearly three times more likely to report shorter times to market for new products and services, and were over twice as likely to report improvement in customer satisfaction, profits, and operational efficiency.
But to get maximum value out of data and analytics, companies need to have a data-driven culture permeating the entire organization, one in which every business unit gets full access to the data it needs in the way it needs it.
This is called data democratization. Doing it right requires thoughtful data collection, careful selection of a data platform that allows holistic and secure access to the data, and training and empowering employees to have a data-first mindset. Security and compliance risks also loom.
Before choosing a platform for sharing data, an organization needs to understand what data it already has and strip it of errors and duplicates.
A big part of preparing data to be shared is an exercise in data normalization, says Juan Orlandini, chief architect and distinguished engineer at Insight Enterprises.
Data formats and data architectures are often inconsistent, and data might even be incomplete. “All of a sudden, you’re trying to give this data to somebody who’s not a data person,” he says, “and it’s really easy for them to draw erroneous or misleading insights from that data.”
Organizations often turn to outside help with data normalization because, if done incorrectly, a business might still be left with data quality issues and can’t get as much use out of their data as intended.
As more companies use the cloud and cloud-native development, normalizing data has become more complicated.
“It might be in a NoSQL database, a graph database, or in all these other types of databases now available, and making those consistent becomes really challenging,” Orlandini says.
In many cases, only IT has access to data and data intelligence tools in organizations that don’t practice data democratization. So in order to make data accessible to all, new tools and technologies are required.
Of course, cost is a big consideration, says Orlandini, as well as deciding where to host the data, and having it available in a fiscally responsible way. An organization might also question if the data should be maintained on-premises due to security concerns in the public cloud. But Kevin Young, senior data and analytics consultant at consulting firm SPR, says organizations can first share data by creating a data lake like Amazon S3 or Google Cloud Storage. “Members across the organization can add their data to the lake for all departments to consume,” says Young. But without proper care, a data lake can end up disorganized and cluttered with unusable data. Most organizations don’t end up with data lakes, says Orlandini. “They have data swamps,” he says.
But data lakes aren’t the only option for creating a centralized data repository.
Another is through a data fabric, an architecture and set of data services that provide a unified view of an organization’s data, and enable integration from various sources on-premises, in the cloud and on edge devices.
A data fabric allows datasets to be combined, without the need to make copies, and can make silos less likely.
There are many data fabric software vendors, like IBM Cloud Pak for Data and SAP Data Intelligence, which were both named leaders in Forrester’s Enterprise Data Fabric Q2 2022 report. But with many available options, it can be difficult to know which to choose.
The most important thing is to analyze and monitor data, says Amaresh Tripathy, global analytics leader at professional services firm Genpact.


