Skip to content
7wData Data and AI tools, companies, events, podcast
  • Tools
  • Companies
  • Podcast
  • Articles
  • Events
  • Newsletter
  • Research
  • Sponsor

Table of Contents

Artificial Intelligence 2021 • By Yves Mulkers

Data lake storage: Cloud vs on-premise data lakes

Data lake storage: Cloud vs on premise data lakes
4 min read
Comma-separated values, Computer file, Data compression
Curated from computerweekly.com →

Handling large amounts of data is a prerequisite of digital transformation, and key to this are the concepts of data lakes and data warehouses, as well as data hubs and data marts.

In this article, we’ll start at the top of that hierarchy and look at data lakes. As organisations try to get a grip of their data and to wring as much value from it as they can, the data lake is a core concept.

It’s an area of data management and analysis that depends on storage – sometimes lots of it – and it’s an activity that’s ripe for a move to the cloud, but can also be handled on-premise.

We’ll also look at the type of storage needed for a data lake – often object storage – and the pros and cons of building in-house or using the cloud.

The data lake is conceived of as the first place an organisation’s data flows to. It is the repository for all data collected from the organisation’s operations, where it will reside in a more or less raw format. Perhaps there will be some metadata tagging to facilitate searches of data elements, but it is intended that access to data in the data lake will be by specialists such as data scientists and those that develop touchpoints downstream of the lake. Downstream is appropriate because the data lake is seen, like a real lake, as something into which all data sources flow, and they are potentially, many, varied and unprocessed. From the lake, data would go downstream to the data warehouse, which is taken to imply something more processed, packaged and ready for consumption. While the data lake contains multiple stores of data, in formats not easily accessible or readable by the vast majority of employees – unstructured, semi-structured and structured – the data warehouse is made up of structured data in databases to which applications and employees are afforded access. A data mart or hub may allow for data that is even more easily consumed by departments. So, a data lake holds large quantities of data in its original form. Unlike queries to the data warehouse or mart, to interrogate the data lake requires a schema-on-read approach.

Sources of data in a data lake will include all data from an organisation or one of its divisions. It might include structured data from relational databases, semi-structured data such as CSV and log files as well as data in XML and JSON formats, unstructured data like emails, documents and PDFs, as well as and binary data, such as images, audio and video. In terms of storage protocol that means it will need to store data that originated in file, block and object storage. But, of those, object storage is a common choice of protocol for the data lake itself. Don’t forget, access will not be to the data itself, but to the metadata headers that describe the data, which could be attached to anything from a database to a photo. Detailed querying of the data often happens elsewhere, not in the data lake. Object storage is very well-suited to storing vast amounts of data, as unstructured data. That is, you can’t query it like you can a database in block storage, but you can store multiple object types in a large flat structure and find out what’s there. Object storage is generally not designed for high performance, and that’s fine for data lake use cases where queries are more complex to construct and process than in a relational database in a data warehouse. But that’s fine because much querying at the data lake stage will be to provide more easily queryable data stores for the downstream data warehouse.

All the usual on-premise vs cloud arguments apply to data lake operations. On-prem data lake deployment has to take account of space and power requirements, design, hardware and software procurement, management, the skills to run it and ongoing costs in all these areas. Outsourcing the data lake to the cloud has the advantage of offloading the capital expenditure (capex) costs of infrastructure to an operational expenditure (opex) one of payments to the cloud provider.

In the 7wData directory

Get the AI & data signal, daily.

335k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

Compare the tools & companies behind this topic

Browse the directory →
  • ChatGPTTool
  • Scikit-learnCompany
  • OONNX RuntimeTool
  • PyTorchCompany
  • TensorFlowCompany
  • PycaretCompany
  • MMxnetCompany
  • SSonnetCompany

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at computerweekly.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.

Want the structural read on any AI or data company?
INS7GHTS

Want a sharper read on this topic?

Ask ins7ghts how the players compare, what people are actually shipping with, and where the trade-offs land.

Tweet LinkedIn Bluesky Threads Email

Related Articles

The Role of Artificial Intelligence in Ethical Hacking
Artificial Intelligence

The Role of Artificial Intelligence in Ethical Hacking

3 min read • 2020
Nine ways machine learning can improve supply chain management
Artificial Intelligence

Nine ways machine learning can improve supply chain management

3 min read • 2020
Facebook’s AI Is Learning to Predict and Prevent Suicide
Apache Hadoop

Facebook’s AI Is Learning to Predict and Prevent Suicide

3 min read • 2017
7wData

Independent reporting on AI and data: daily newsletter, podcast, deep dives.

Read

  • Ins7ghts newsletter
  • AI Beat newsletter
  • Latest articles
  • Podcast
  • Research guides

Use

  • Tools directory
  • Company directory
  • Research
  • Events
  • ins7ghts

Company

  • About
  • Contact
  • Sponsor a slot
  • Media kit
  • RSS feed

Follow

  • LinkedIn
  • X
  • YouTube
  • Instagram

© 2026 7wData. Independent. Belgium-based.

Privacy Cookies Terms Imprint Cookie settings
INS7GHTS
New · ins7ghts Drops

The AI governance conversation already moved. Most 2026 plans missed it.

Drop #1 · 60 pages · Launch week €99 (then €149) · ends Thu 9 July

Read Drop #1 →
Cookies on 7wData

We use strictly necessary cookies for the site to work, and optional analytics cookies to understand how readers use 7wData. We never share your data with advertisers. See our Cookie Policy.

Get the AI & data signal, daily. 335k+ already do.
Thanks. Check your inbox to confirm.
Get the AI & data signal

One curated email a day. 335k+ data & AI professionals already read it.

No spam. Unsubscribe anytime.

Check your inbox.

We just sent a confirmation. Click the link to start receiving the daily signal.