Scaling Your Data Strategy

Companies growing at a fast pace enjoy two unique advantages simultaneously. They are able to utilize small but mighty teams with a can-do attitude and deliver product features quickly to the market. At the same time, in order to maintain rapid expansion, they have a tangible incentive to pay just as much attention to setting up the fundamental practices for sustainable growth. This means managing digital transformation to take advantage of the benefits of new technology, and managing increasing simultaneous and interconnected projects. Not to mention that at the intersection of all these projects and processes, there is often a lot of data zipping around! As data grows exponentially, this unique position can be leveraged to craft a robust, long term data strategy.
We attempt to address large-scale problems with data-driven solutions. In this context, business relationships can quickly become complex, and identifying patterns and behaviors around your data can become incrementally challenging. If you are regularly starting new projects that are deeply intertwined with existing features and processes at your organization, you are no longer able to build isolated, one-off built-from-scratch solutions. Your operation has grown to the point that it pays off to have a corporate data strategy. In this article, I will present my vision for a cohesive data strategy based on my prior professional experiences.
No matter where you are in your data-driven journey, having a data strategy helps unlock the power of data and allows your organization to treat data as a critical asset. A data strategy is a plan designed to improve the methods, practices, and processes of data used across the organization, and to ensure that data is being used in a sustainable and reproducible way. Because data is generated and used by diverse business units with different practices and responsibilities, having a committee to oversee a data strategy across the organization becomes central to business success. Not having a data strategy is okay depending on where you are in your journey, but it becomes more likely that different departments will solve data issues on their own, resulting in wasted resources, units operating in silos, and a growing lack of cohesiveness in the organization. I managed to encapsulate that grim possibility in one sentence, so let’s focus on positive and constructive ideas henceforth.
A robust data strategy should generally cover the following core topics:
This involves understanding all the core entities of the organization’s data such as: customers, locations, product, transactions, and their relationships. In the early stages of an operation, it would have been okay to represent these relationships in a relational database schema, but as the various data types, processes, and storage formats evolve; it is important to have a managed data catalog with definitions, meanings, and relationships to link disparate systems together.
Newer methods of combining data under a single interface such as a data lake, help to avoid a lot of the problems and complexity of managing and using data in organizations. Even so, consolidating the terminology, metadata, and relationships under one catalog makes accessing and using data much simpler. For each field of data, such a catalog might include definitions, sources, locations, domains, use-cases, stakeholders, etc.
Often the idea of data governance seems restricted to users and the analytics environment, but in reality, we are using and generating data in every operation. Once data is decoupled from the application that created it, the organization should define rules of and details for the data so that all stakeholders can understand how to make use of it.


