Do Customers Want Open Data Platforms?

4 min read
Curated from datanami.com →

Snowflake turned some heads in the big data market with recent blog posts and articles that cast doubt on the benefits of open data architectures. Customers should “choose open wisely,” the data warehouse giant says, and avoid “fall[ing] into the trap of letting the means get confused with the end.”

Snowflake lobbed the opening salvo in its latest battle against open platforms with a 3,800-word blog post on March 31 written by founders Benoit Dageville and Thierry Cruanes, along with two other executives. Dageville followed that up with a 2,300-word opinion piece in Infoworld’s Insider, which is not open to the public and requires reader registration (although the first taste is free).

In summary, Snowflake cautions against embracing open in all its permutations–open data standards, open format, and open source software–without careful consideration of the possible downsides.

“We see strong opinions for and against open, we see table pounding demanding open and chest pounding extolling open, often without much reflection on benefits versus downsides for the customers they serve,” Snowflake execs write in the blog piece. “For many organizations who have fallen into the trap of assuming that open is synonymous with innovative and cost-effective, they have learned the hard way that neither is the case.”

To be sure, there have been some colossal implementation failures with open data platforms. It wasn’t long ago that the big data ecosystem was dominated by Hadoop, which many saw as the open answer to long closed data warehouse platforms. But the stunning fall of Hadoop, along with the meteoric rise of cloud data lakes and warehouses–Snowflake included–shows how quickly ideas about data architecture can change.

Snowflake was an early detractor of Hadoop. Former CEO Bob Muglia had an especially harsh take on the failures of that open data platform. Under new CEO Frank Slootman, Snowflake has found tremendous success with its cloud data warehouse. This success can be traced, at least in part, to Snowflake offering much simpler user experience for analytics customers, at least compared to setting up and running a Hadoop-based system.

But how much of Snowflake’s success is tied to the fact that it stores and processes data in a proprietary data format? That’s an open question, but one that the company clearly felt compelled to answer. In the blog post, the company makes a compelling case that shielding user from technical complexity is worth giving up that control.

“At first glance, the idea of any data consumer or any application being able to directly access files in a standard, well-known format sounds appealing,” the Snowflake executive write. “Of course that is until a) the format needs to evolve, b) the data needs to be secured and governed, c) the data requires integrity and consistency, and/or d) the performance of the system needs to improve.

“What about an enhancement in the file format that enables better compression or better processing?” they ask. “How do we coordinate across all possible users and applications to understand the new format? Or what about a new security capability where data access depends on a broader context? How do we roll out a new privacy capability that reasons through a broader semantic understanding of the data to avoid re-identification of individuals?”

The challenges of pursuing transactional integrity, performance optimization, and application coordination further add to the misery of the open data architect, they add.

“Decades of experience navigating through these very trade-offs give us a strong understanding of and conviction about the superior value of providing abstraction and indirection versus exposing raw files and file formats,” the executives write. “We strongly believe in API-driven access to data, in higher level constructs abstracting away physical storage details. It’s not about rejecting open; it is about delivering better value for our customers.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at datanami.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.