Excuse Me, Do You Speak Data?

6 min read

Just like many other analytics professionals I regularly speak to, I sometimes find it hard to explain what I exactly do in my day to day job to people who don’t work in the field despite the big hype in the media around anything data in the past few years (e.g. big data, data science, etc.). While it might not be the end of the world if your friends don’t understand what you do at work the consequences of such misunderstandings can be a lot more significant if the same situation happens in your workplace with stakeholders with whom you interact frequently without them fully understanding the significance of your work and how you can bring more value to them.

A couple of weeks ago while reading the book “Infonomics” written by Douglas B. Laney, I came across a new concept called Information as a Second Language (ISL) that was developed by a Gartner analyst named Valerie Logan. The concept was modelled after the teaching of English as a Second Language (ESL) which is very common in anglophone countries such as the UK, Canada, and the US that attract international students, workforce, and immigrants who need to immerse themselves in the local language in order to be able to function in their new environment. The concept of ISL is about the attempt to standardize existing data-related dialects used by employees of different business and IT functions since poor data literacy and miscommunication came up as the second biggest barrier for the execution of data and analytics strategies among chief data officers (CDOs) surveyed by Gartner in 2018. As an avid learner of languages, especially Romance languages, I realized there’s one country that could potentially give us a pretty good case study about the ways we could solve the problem of data literacy by creating a standardized data jargon for the organization and what we should realistically expect as an outcome. The country I am going to talk about is Italy.

Up until the year of 1861 when Italy was unified as a country as well as the modern Italian language was standardized, the inhabitants of the Italian peninsula lived in small states and spoke different dialects that evolved from older forms of the spoken Latin language that date back to the days of the Roman Empire. This situation was unique in Europe back in those days. Yet, within less than 150 years Italy was able to reach fluency of the new standardized language among the overwhelming majority of its 60 million citizens despite a high rate of illiteracy before the year 1861 (78% of the population). Even if many Italian dialects are still used today in Italy’s different regions (Sicily, Veneto, Calabria, etc.) mainly for everyday spoken communication, the standard language is widely used anywhere around the country and most people who do business in Italy use the modern language.  Here’s what we could learn from the case of Italy and might be able to apply in organizations as part of a solution to overcome the data literacy barrier (hopefully within a lot less than 150 years):

Role models – often, it is easy to introduce a new way to use a terminology or jargon when enough role models who’ve already mastered it already exist in an environment. In the case of Italy, it was the Tuscan dialect spoken in Florence and Tuscany that gained popularity in the 1300s onward thanks to writers such as Dante Alighieri who made this dialect visible throughout the Italian peninsula and desirable among writers. Later on, the Tuscan dialect served as the foundation for the modern Italian language due to its high exposure. If we learn from this example we need to think of areas of the business (and I guess it can vary by industry) that are the most data-savvy and can help identify the dominant “data dialect” that can serve as a foundation for the standard data jargon in the organization. From there, it is a question of finding someone who mastered the language and can help take it to the next level by actively working on glossaries and terminology.

Assessing the gaps – this step applies to all languages. When a learner of any language wants to enroll in a class, he or she must go through an assessment of their level in order to be classified as a beginner, intermediate, or advanced. Italy needed to increase the number of literate people and move from a situation of having 78% of its population being illiterate prior to 1861 in order to increase the number of people who can master the unified Italian language. In the case of Italy, the Italian constitution of 1948 gave everyone the right to basic education. What could you do in your organization? In order to assess the current literacy rate and the initial diffusion potential of a cross-organizational data jargon, it might be a good idea to assess the general understanding of basic information-related concepts among business users and understand where the organization stands as a whole, and what terminology is already used but barely understood by most employees. If the overall diffusion potential is low since not many people are tech-savvy or need to make data-driven decisions daily, it might be better to start with teaching general information concepts before jumping into the development of an organizational jargon around data. Many organizations are still figuring out an initial data and analytics strategy and might need this initial step first.

Means of diffusion – In the 1950s and the 1960s, the radio and the TV helped diffusing the Italian language across the country even further and played a role in increasing the level of literacy among the entire population, especially in lower-income areas where people were less educated. An organization needs to find a platform that could help with the diffusion of a new unified data jargon, such as a series of lunch-and-learn events, workshops, and even a centralized portal that includes videos or other educational aids in order to serve as means of the diffusion of a standardized way to communicate information and data.

Expectation management – it’s important to remember that despite the good results we saw in Italy in the adoption of a new language, many of the dialects still exist and are used in daily life in different regions. It is very likely that data dialects will not completely disappear in organizations over time due to different metrics monitored in different areas of the business, but it is possible with the right strategy to develop a lingua franca that will connect the different areas of the business that use those dialects with more general concepts and terminology that could be used across the organization.

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.