3 Challenges of Storing Billions of Files Beyond Big Data

3 min read
Curated from inside.igneous.io →

Enterprises collect more data than ever but life at the speed of data presents a number of challenges. Igneous recently delved into the topic of challenges of massive file systems that go beyond what is generally referred to as “big data” by bringing together a group of people on the infrastructure frontlines to get their insights.

Thanks to everyone who participated in our recentCrowdChat conversationon massive file system data. Led by host John Furrier, participants explored the data management and backup challenges that enterprises experience when Networked Attached Storage (NAS) reaches billions of files. Plans are in the works for a follow-up CrowdChat conversation and we’ll let you know when we schedule it.

Before I share the insights from our Feb. 7 CrowdChat, let’s look at a few statistics. Exactly how much data are we talking about, anyway?

Big data is just the start.

Here are three big challenges faced by enterprises generating massive amounts of data:

Traditional backup is a “necessary evil” in most enterprises because long-term snaps are pricey and don’t protect data in many scenarios, according to Jeff DiNisco, P1 Technologies’ Vice President of Solutions Architecture.

Another challenge when developing a solution for backing up your data? The way that enterprise executives view data storage, according to John, the CrowdChat host’s CEO. While he considers backup “super critical,” others don’t always see the value.

“Most organizations see data storage as a cost center rather than a potential gold mine,” he said.

Get the AI & data signal, daily.

335k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

Deciding what data to back up depends on the governing requirements regulating the user and vertical, according to Webair Chief Technology Officer Sagi Brody. The Health Insurance Portability and Accountability Act of 1996 is one example.

“Some, like HIPAA, don’t have an exact de facto standard and are up for interpretation,” Sagi said.

Manual vs. automatic. Involved users vs. uninvolved users. CrowdChat participants disagreed on the keys to successful migration data among tiers.

The success of tiering depends on the clarity of the data lifecycle, according to P1 Technologies’  Jeff DiNisco.

“Tiering works when the lifecycle is clear,” Jeff said. “When it’s not, it can end up pointless, especially when there’s little cost differentiation.”

Public content networks such as Netflix and Comcast turn to tiering as a solution to large, unstructured data, said Webair’s Sagi Brody. With the ever-increasing size and complexity of big data and massive data, and the costs of private line/bandwidth lag, “enterprises will be forced to do this,” Sagi said.

Bryan Champagne “completely” agrees. Bryan is the Co-founder of Congruity, which was formed in the merger of MSDI, Source Support Services, and Rockland IT Solutions.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at inside.igneous.io →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.