Skip to content
7wData Data and AI tools, companies, events, podcast
  • Tools
  • Companies
  • Podcast
  • Articles
  • Events
  • Newsletter
  • Research
  • Sponsor

Table of Contents

Big Data 2017 • By Yves Mulkers

If Big Data Is the New Crude, Data Virtualization Is the New Refinery

If Big Data Is the New Crude
3 min read
Big Data, Clickstream, competitive intelligence
Curated from datacenterjournal.com →

Big data is like an abundant, expanding natural resource emerging from the modern data landscape. IoT (sensor), mobile, social, clickstream, web and open data are important contributors to the proliferation of data we’re witnessing today. Worldwide data is expected to increase tenfold by 2025—reaching a total of 163 ZB—according to a recent IDC-Seagate study.

Data is plentiful, but not necessarily useful in its raw, unrefined form. As with any natural resource, “crude” data must be refined before it can be harnessed for productive purposes, such as equipment maintenance, product innovation, competitive intelligence, marketing, data monetization and active health care. The refinement process can incorporate data exploration, preparation, correlation and contextualization, labeling and annotating, unification and integration, and application of security and governance policies. Metadata is also an important component, as it serves a role in both the input and output stages of the overall data-refinement process.

The extent to which data analysis contributes to unbiased conclusions, accurate predictions and insightful decision-making is constrained by the veracity of that data. If it hasn’t been provisioned for analysis, the data may suffer from fragmentation, minimal labeling and missing information. Such characteristics can be evident in electronic health records (EHRs), which illustrate the challenges of data refinement. One hurdle to gathering and analyzing EHR data is the scarcity of proper labeling and consistent semantics.

EHRs are designed primarily to fulfill patient-care, administrative and financial needs. The multipurpose objectives of EHRs—which don’t take into account data analysis per se—can create data fragmentation, which requires rectification before the data can be provisioned for analyses such as clinical research. Another challenge to building data sets from shared patient health records is the lack of standardization in how EHRs are implemented among health-care organizations, and even within the same health-care system. For example, distinct departments (e.g., radiology, orthopedics and internal medicine) in the same hospital may employ EHRs differently to satisfy their unique data-entry requirements, documentation and ordering needs, and preferences, thereby creating data silos.

Data security and privacy can also be impediments to analyzing regulated data, such as that in EHRs. The best approach to surmounting this obstacle is applying proper security and governance during the refinement process. Companies such as Google are experimenting with federated learning in their effort to advance analytics while ensuring privacy.

Data refinement is crucial to achieving reliable outcomes from data analysis, including meaningful conclusions, accurate predictions and informed decisions. Ideally, the process of refining raw data to produce complete and meaningful information does the following:

Modern analytics relies on data from myriad fragmented data sources. Experience tells us that big data sources aren’t always amenable to replicating and relocating when the data is distributed across multiple systems. Data virtualization delivers the scale to work effectively with big data sources by offering an alternative paradigm: move processing to the data.

In the 7wData directory

Get the AI & data signal, daily.

335k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

Compare the tools & companies behind this topic

Browse the directory →
  • LakebaseTool
  • SurrealDBCompany
  • WeaviateCompany
  • ClickHouseTool
  • MySQLCompany
  • Astra DBTool
  • OceanbaseCompany
  • Data TransferTool

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at datacenterjournal.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.

Want the structural read on any AI or data company?
INS7GHTS

Want a sharper read on this topic?

Ask ins7ghts how the players compare, what people are actually shipping with, and where the trade-offs land.

Tweet LinkedIn Bluesky Threads Email

Related Articles

How data mining helps in business intelligence
Artificial Intelligence

How data mining helps in business intelligence

3 min read • 2022
How artificial intelligence and machine learning are changing the development landscape [Q&A]
Artificial Intelligence

How artificial intelligence and machine learning are changing the development landscape [Q&A]

3 min read • 2022
5 Top Data Challenges That Are Changing The Face Of Data Centers
Big Data

5 top data challenges that are changing the face of data centers

4 min read • 2017
7wData

Independent reporting on AI and data: daily newsletter, podcast, deep dives.

Read

  • Ins7ghts newsletter
  • AI Beat newsletter
  • Latest articles
  • Podcast
  • Research guides

Use

  • Tools directory
  • Company directory
  • Research
  • Events
  • ins7ghts

Company

  • About
  • Contact
  • Sponsor a slot
  • Media kit
  • RSS feed

Follow

  • LinkedIn
  • X
  • YouTube
  • Instagram

© 2026 7wData. Independent. Belgium-based.

Privacy Cookies Terms Imprint Cookie settings
INS7GHTS
New · ins7ghts Drops

The AI governance conversation already moved. Most 2026 plans missed it.

Drop #1 · 60 pages · Launch week €99 (then €149) · ends Thu 9 July

Read Drop #1 →
Cookies on 7wData

We use strictly necessary cookies for the site to work, and optional analytics cookies to understand how readers use 7wData. We never share your data with advertisers. See our Cookie Policy.

Get the AI & data signal, daily. 335k+ already do.
Thanks. Check your inbox to confirm.
Get the AI & data signal

One curated email a day. 335k+ data & AI professionals already read it.

No spam. Unsubscribe anytime.

Check your inbox.

We just sent a confirmation. Click the link to start receiving the daily signal.