Best Practices for Setting Up a Data Warehouse in the Cloud

4 min read
Curated from dzone.com →

Data and analytics have become an important part of today’s business. Modern businesses generate large amounts of data from their customers, suppliers, partners, and internal systems. If carefully analyzed and interpreted, this data could provide immense insight into business growth and sustainability. The data collected from these different data sources comes in unstructured, semi-structured, and structured forms. It’s useful to transform the data into a structured format and send it to a central location. This is where businesses rely on data warehousing solutions to work as a central repository, where the data is collected for analytics as well as different tools used for data transformation.

Running a data warehouse is becoming increasingly expensive and difficult to manage and scale. While costs amount to millions of dollars upfront for both software and hardware, it also takes months to plan, architect, and implement the data warehousing solution, which is not always viable for small and medium-sized businesses. There are different types of data warehouse solutions out there which need to be selected depending on the individual business requirements.

Many businesses are therefore moving their data warehouses to the cloud to improve performance and decrease costs. With the cloud, scalability and elasticity are built-in, and it’s possible to expand for rapid data growth, both in terms of processing capacity and storage.

Since the cloud provides various tools, managed services, and technologies to reduce the complexity and management overhead of the data warehousing solutions, it is possible for the business to focus on using the data to generate results rather working on the technology and managing the data warehouse itself. Let’s look at several best practices in using the cloud for data warehousing and the advantages it provides.

When moving to the cloud for a data warehousing solution, it is required to migrate data from existing solutions to the cloud. Depending on the migration strategy, it is possible to also move part of the data pipeline to the cloud, in addition to moving structured data from the existing data warehouse. This could even include moving unstructured or semi-structured data to the cloud to store and transform the data, as required by the data warehousing solution.

For example, instead of maintaining a file server locally, it is possible to directly ingest unstructured data such as CSV or Excel files to an object storage service such as Amazon S3. The data could then be transformed and processed using a data pipeline within the cloud for increased performance, reduced costs, and reduced management overhead.

In addition, depending on the data volumes and operational requirements, it is required to set up the migration strategy to move data to the cloud data warehousing solution. For example, if the data volume is low (which doesn’t require a continuous operation), it is possible to set up a one-step migration to extract the data and import it into the cloud data warehousing solution.

However, when dealing with large data volumes with continuous operational requirements, this approach becomes impractical due to the massive amount of data to be moved as well as the new data being added within the migration time frame. Therefore, this requires a two-step process. for the first step, it is important to first extract the data in non-peak usage to minimize the impact to the existing data warehouse and migrate to the cloud using a medium that matches with the extracted data volume. For example, if the data volume is large, and expands to a couple of terabytes or petabytes, it is viable to physically store the data in a storage device and shift it to the cloud, rather than sending it across the wire. There are services such as Amazon Snowball used to simplify the process by leasing the required storage devices.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at dzone.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.