Clarifying Hybrid Cloud Storage: Passive vs. Active Data

I recently gave a webinar with Forrester analyst, Richard Fichera. In his opening, he noted “the inexorable march to public cloud is underway.” And he’s right. We see it everywhere. But “march” is a critical word. It’s not a leap and it won’t happen overnight. Along the journey, many enterprises will implement hybrid cloud first.
As such, we see hybrid cloud increasing in popularity across the industry. For example, hybrid cloud adoption tripled in 2016, climbing to 57% from 19% of organizations surveyed, according to an Intel Security report.
Yet as hybrid cloud adoption widens, I’ve noticed that often there is still confusion on what it means for different parts of the IT infrastructure stack. So I thought it would be useful to walk through what we mean by hybrid cloud as it relates to storage, and software-defined storage (SDS) in particular.
Basically, when it comes to SDS products, there are two scenarios:
Let’s discuss the pros and cons of each.
The first scenario is where, in fact, the bulk of the market seems to be today. Essentially, there is a “southbound” object storage interface whereby the storage software puts the object into a public cloud such as AWS S3 or AWS Glacier. In this instance, the storage solution acts like a cloud gateway and can get and put data in a public cloud storage “bucket.” The pros of this approach are that it’s simple, extensible, and cost effective. It helps eliminate the need for a disparate cloud gateway solution and treats the cloud like a cold storage tier. Think of it as the modern equivalent of tapes.
However, the cons of this scenario are a bit more complicated.
In this scenario, the data that you’ve stored in the cloud is not “active.” Because you’re storing just a copy of that data in the cloud, to do anything with it, you need to restore it first. In effect, the problem you’re solving here is one of disaster recovery (DR). Certainly, this makes sense in backup and archive use cases where you want multiple copies of that data should your data center go down.
As such, getting the data back so you can do something with it (e.g., run analytics to surface actionable insights) can be time-consuming and will most certainly incur cloud fees. Based on the cloud provider you’re using, there are considerable retrieval time and cost considerations associated with this data recall.
The vast majority of storage solutions boasting hybrid cloud capabilities fall into this first scenario. Essentially, it’s attaching to the cloud as though the cloud were another storage tier inside your on-premises storage solution, usually along the lines of tiering from flash to spinning disk and then to the public cloud.
There’s nothing wrong with the approach in this scenario, but it’s not how we think about hybrid cloud storage at Hedvig.
Using us at Hedvig as an example, when we talk hybrid cloud storage, what we mean is that you actually run a full-fledged instance of the Hedvig Distributed Storage Platform in the public cloud. The difference, in our case, is that Hedvig is running on the compute side of the public cloud, such as Amazon EC2. We take these cloud servers running our software, virtualize and aggregate disk capacity (from a service like Amazon EBS, for example), and join these as active nodes in a hybrid cloud cluster. Here the data is active — it’s fully available to whatever application needs it and can be accessed locally in that public cloud.
Because the data is fully available, you can treat the public cloud as a full data center environment.


