Best Practices for Managing Data both Big and Small

In 2013, an electronic music band from Brooklyn appeared on the music scene with the name Big Data. Clearly, the lexicon of “big data” has become a household phrase over the last decade. Yet, what does big data really mean? How is it managed differently from other data? And is more data always better?
In simplistic terms, big data is just more; more complexity, more volume, and more velocity. One example of this pace of change is that ten years ago I tracked my steps per day with a pedometer. Now I can wear a Fitbit that measures my number of steps, heart rate, stairs climbed, sleep patterns, and more for every hour of the day. With a Fitbit my data are more complex, the volume of the data is larger, and the velocity of information is higher. Wearable technology is just one example of the application of big data.
Yes and No. Data management best practices are the same regardless of the size of the data. However, the need to organize data efficiently is magnified as the size and complexity grow. For a sound data management program there needs to be a focus on the fundamentals before you can have reliable data analytics and business intelligence. In Maslow’s hierarchy of needs you need food and safety before you can have love, esteem, or self-actualization. Focusing on love or esteem when you don’t have food or safety makes it hard to survive. Similarly, trying to create an analytics program without an appropriate technical and governance infrastructure is, at best, inefficient.
It is common for an analyst to start with data management by getting some data and some statistical software. Under this approach, the analyst may spend 80% of their time doing prep work to organize and understand the data and 20 percent performing analysis. This approach works for an isolated project but does not enable or support an enterprise analytics program.
The most often overlooked fundamental of data management is the curation of metadata. Metadata is the information that describes the data–much like the label on a soup can. Metadata tells you what you have and what it means and there are several kinds of metadata. Metadata, can be a catalog, much like a library card catalog tells you about the library’s book collection, without giving you access to the data. Metadata can also be organized at data element level to give meaning to the data items within the dataset much like the information on the soup can label tells you about the contents.
As complexity, volume, and velocity increase, the need for adherence to best practices becomes critical
In fact, I often equate data management to a can of soup. The soup is the data, the can is the database, and the label is the metadata. If you rip the label off the can, you can still eat the soup but you may not know if it is cream of mushroom or cream of celery soup until you open the can. Now imagine a pantry full of cans with no labels. To understand what you have you would need to open all of the cans and examine the contents. When the only metadata maintained about the content of a file or dataset is the name of file you have the equivalent of a pantry where the only labels are soup, beans, vegetable, fruit, and other. It makes cooking challenging, and if you have to worry about food allergies, potentially dangerous.
Another data management fundamental is the use of data standards. In fact, standardized definitions are often overlooked or undervalued. However, standards ensure that (things) are interoperable, easier to work with, and more efficient. Our lives are filled with standards, whether it’s standards that apply to our electrical outlets, internet protocols, or the dictionaries we use.


