The hardest parts of cloud data management

Today’s organizations store and manage their applications and data across heterogeneous environments — from on-premises data centers to edge to public cloud environments. However, many often find themselves challenged with managing duplicate copies of data, unpredictable or uncontrollable costs, differences in storage performance, and capabilities across environments, and a lack of clear data governance and controls. With applications and services spread across environments and without a clear data integration approach, data management becomes incredibly complex. For example, data is commonly transferred and loaded into multiple different environments, proliferating copies of data. These copies of data both widen governance and security challenges, but also exacerbate new cost challenges. As traditional IT organizations increase their in-cloud data footprint, they are often caught off-guard by unexpected and hard-to-predict forecast costs surrounding the usage of their data. While many are accustomed to thinking about costs of storage, costs of data access (e.g., transit and access charges, added fees for burst performance, etc.) are not typically encountered on-premises and require operational changes to manage. But cost control is not the only area often requiring operational or application changes — customers must often manage differences in data service availability, capability, reliability, and available performance levels. Cloud infrastructure and platform services (IaaS, PaaS) offer the freedom and flexibility of the cloud operating model to IT organizations. Realizing those benefits in-cloud versus on-premises, however, requires awareness and diligent management of these challenges.
Poor data quality and data integration issues, coupled with a lack of data discoverability, are often some of the biggest challenges facing organizations today when it comes to managing data in the cloud. According to recent research from MIT Technology Review and Databricks, data leaders reported that their teams spend 41% of their time on data integration and preparation, and nearly every respondent (96%) reported negative business effects as a result of data integration challenges. Data can be a powerful tool, but if organizations are spending a disproportionate amount of their time cleaning, organizing and migrating their data instead of analyzing and taking action from it, the value is lost. Moreover, the effects can significantly impact the bottom line – with a lack of discoverablity for the rightdata, teams face massive data duplication issues and are far less productive with need for more manual data scrubbing. With this, decision-makers may unknowingly be relying on incomplete or inaccurate data. Adding to this, many of today’s enterprises are adopting a multi-cloud approach, wanting the efficiency of cloud scalability without being locked-in with a single provider. This multi-cloud approach offers greater flexibility and resiliency in a data strategy but is not without challenges as data teams may need to reimplement workloads and data models between platforms and, ultimately, need the convenience of leveraging common data tools that can work seamlessly across each cloud.
Data is moving to the cloud because it is an excellent place to store, manage, and analyze data. The cloud breaks down information silos that exist in on-premises computing, making it much easier to share data internally and with business partners and customers. However, when you put all your data in one place, you also must implement safeguards that govern the use of the data — most importantly data access control. This has proven to be a challenge for technology vendors and for the organizations that are managing their data in the cloud. The underlying problem is caused by SQL. The industry-standard database query language is a core element of the Modern Data Stack, which is the ecosystem of technologies that enable us to manage data in the cloud.


