Who should be responsible for your data? The knowledge scientist

How can you build a data-driven culture and spur digital transformation without thinking through who should be responsible for your data? Let’s do that together.
Data engineers and data scientists each occupy critical roles. Data engineers manage the data infrastructure and are in charge of designing, building, and integrating data workflows, pipelines, and the ETL process. Their goal is to provide data for data scientists’ analysis. Data scientists are those who can turn data into insights by applying statistics, machine learning, and analytical approaches. Their goal is to answer critical business questions.
Data-driven organizations require reliable, clean data to function. Without it, your AI, machine learning, and analytics are worthless. Unreliable, erroneous, and incomplete data leads to answers that can’t be trusted—hence, “garbage in, garbage out.”
Therefore, the process of wrangling and cleaning data is crucial, often said to be 80% of a data scientist’s work. Typically, this is seen as boring, annoying grunt work people don’t want to do.
However, I think this negative view is at least partly based on a major underappreciation of the significance of such work. Data wrangling and cleaning is not simply about eliminating white spaces, replacing wrong characters, and normalizing dates. Stepping back, these tasks should be viewed in the context of two key objectives:
Yes, data wrangling and cleaning can take 80% of a data scientist’s time and energy. This does not mean that 80% is wasted. While these tasks can and should be optimized for efficiency, they are part of the vital knowledge work that should be elevated within a data-driven organization. But who should be doing it?
In typical organizations, the need for reliable data is constant, but the knowledge work that creates it is ad hoc. Practices and results are not documented and shared because data scientists are usually not equipped, trained, or incentivized to do so. Indeed, in our experience, a lot of the “softer” knowledge work (like conference calls, discussions, whiteboarding sessions, documentation, long Slack chats) required to create clean and reliable data is not valued by data scientists or their managers. Making matters worse, most tools are designed and provisioned for a small set of user types and teams to the exclusion of other user types and teams. Thus, the responsibility to create and manage reliable data is siloed, scattered, or even non-existent.


