Skip to content
7wData Data and AI tools, companies, events, podcast
  • Tools
  • Companies
  • Podcast
  • Articles
  • Events
  • Newsletter
  • Research
  • Sponsor

Table of Contents

Data Architecture 2022 • By Yves Mulkers

why is effective data collaboration so elusive?

Why Is Effective Data Collaboration So Elusive?
4 min read
Apache Hadoop, API, Conceptual model
Curated from diginomica.com →

Traditionally, the job of gathering and integrating data for analytics fell on data warehouses. Data warehouses cleaned and aggregated the “operational exhaust” of business transactions.

Data warehouses created precision, accuracy and consistency out of messy transactional structures, so that business people could look back, and make sense of the past.

As data warehousing evolved, so did our operational applications. With the advent of the Internet, a whole new challenge to precision, accuracy and consistency emerged.

For business analysts and data scientists alike, their work now involves much more searching and discovery of data rather than navigating through steady structures. It is no longer a singular effort. Sharing results and persisting new views of data and models, and collaboratively governing their catalogs is the approach that works.

It is important to make the distinction between tools that provide “connectors” that operate on a physical level, either generating SQL from structural metadata or using API’s, etc., and those that have a rich understanding of the data based on its content, structure and how it is being used. Connectors are not sufficient. Someone, or some thing, has to understand the meaning of the data to provide a catalog for analysts to do their work. 

Keep in mind that all of the data used by data scientists, analysts and other applications is essentially “used.” In other words, it was originally created for purposes other than to be analyzed: supporting operations, capturing events, recording transactions. The data at source, even when it is clear, consistent and error-free, which is rare in an integrated context, will still contain semantic errors, missing entries, or inconsistent formatting for the secondary context of analytics. It must be handled before it can be used. This is especially true if it is meant to be integrated with data from other sources or other previously managed data.

Even somewhat stable and understandable data like logs, especially weblogs, tend to drift over time. Can you think of a major website that hasn’t gone through a complete refresh in the past three years or so?

Data warehousing addressed this problem of dealing with “used” data long ago, with processes that executed before data was stored. This is where data warehousing and Hadoop differ. Tools such as ETL, methodologies and best practices ensured that any analyst working with that data accessed a “single source of truth” that was already cleaned and aggregated to produce a pre-defined business metric, often labeled a key performance indicator.

Though in fairness, applications from data warehouses are often quite creative. These solutions are also partially useful for data scientists, but because the data was pre-processed to fit a specific model or schema, the richness of the data is lost when it comes to hypothesis testing in a more exploratory manner and typically too slow to implement when experiments and discoveries are happening in an unplanned fashion (“Our competitor just released a press release about a new pricing model, how should we respond?”). 

The innovation of widespread cloud computing and “cloud native” data warehouses introduced a better storage mechanism and best practices for hypothesis testing on rich, raw data – at low cost. These cloud data warehouses, as well as cloud native storage protocols such as JSON, evolved out of a system designed to capture “digital exhaust.”

Initially the byproduct of online activities, today, “clouds” also include a wealth of machine-generated data; events gathered from sensors in real-time. While data warehouses typically stored aggregated information based on application transactions, digital exhaust can be found in a multitude of forms like XML and JSON, typically referred to as “unstructured,” though more accurately defined as “not highly structured.

In the 7wData directory

Get the AI & data signal, daily.

335k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

Compare the tools & companies behind this topic

Browse the directory →
  • LakebaseTool
  • SurrealDBCompany
  • WeaviateCompany
  • ClickHouseTool
  • MySQLCompany
  • ClouderaCompany
  • Astra DBTool
  • OceanbaseCompany

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at diginomica.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.

Want the structural read on any AI or data company?
INS7GHTS

Want a sharper read on this topic?

Ask ins7ghts how the players compare, what people are actually shipping with, and where the trade-offs land.

Tweet LinkedIn Bluesky Threads Email

Related Articles

Asset Managers Attack Data Silos with Governance
Data Management

Asset Managers Attack Data Silos with Governance

3 min read • 2016
IoT and AI push tech ethics to the forefront of development
Big Data

IoT and AI push tech ethics to the forefront of development

3 min read • 2021
How to start a data analytics project in manufacturing
Big Data

How to start a data analytics project in manufacturing

3 min read • 2017
7wData

Independent reporting on AI and data: daily newsletter, podcast, deep dives.

Read

  • Ins7ghts newsletter
  • AI Beat newsletter
  • Latest articles
  • Podcast
  • Research guides

Use

  • Tools directory
  • Company directory
  • Research
  • Events
  • ins7ghts

Company

  • About
  • Contact
  • Sponsor a slot
  • Media kit
  • RSS feed

Follow

  • LinkedIn
  • X
  • YouTube
  • Instagram

© 2026 7wData. Independent. Belgium-based.

Privacy Cookies Terms Imprint Cookie settings
INS7GHTS
New · ins7ghts Drops

The AI governance conversation already moved. Most 2026 plans missed it.

Drop #1 · 60 pages · Launch week €99 (then €149) · ends Thu 9 July

Read Drop #1 →
Cookies on 7wData

We use strictly necessary cookies for the site to work, and optional analytics cookies to understand how readers use 7wData. We never share your data with advertisers. See our Cookie Policy.

Get the AI & data signal, daily. 335k+ already do.
Thanks. Check your inbox to confirm.
Get the AI & data signal

One curated email a day. 335k+ data & AI professionals already read it.

No spam. Unsubscribe anytime.

Check your inbox.

We just sent a confirmation. Click the link to start receiving the daily signal.