The Data Value Chain: Moving from Production to Impact

The data value chain describes the evolution of data from collection to analysis, dissemination, and the final impact of data on decision making. While the value chain can be applied to all types of data, we hope it will be particularly useful for a better understanding of the gaps in gender data.
Data are the information we use as the basis for reasoning, analysis, and debate. They are the factual currency for evidence-based policy making. In a data-driven utopia, data would be highly valued and demanded and used ethically and effectively. But data travel a long journey, gaining value as they go, before they achieve their highest purpose. The data value chain provides a framework through which to visualize the life cycle of data, from defining a need to using them for impact.
The concept of a value chain was first used by John Porter to describe how a firm receives raw materials, turns them into products, distributes, markets and sells them, and subsequently provides services, adding value and hence, increasing the firm’s revenues with each step. [1] Value chain analysis has since been applied to sequential activities at a larger scale: a global value chain has been used to characterize the movement of goods in various stages of production between firms and across national boundaries. [2]
The value chain describes connections between each step that change low-value inputs into high-value outputs. Although it has a logical flow, from start to finish, a value chain has no theory: it is a pragmatic construct. Value chain analysis can help to identify impediments toward a final goal. It can also focus attention on high-value steps where more effort is needed or suggest ways to reduce resources spent on low-value activities. The model of the value chain is just as applicable to the production and use of intangible goods, such as data and statistics, as it is to physical products.
The data value chain has been discussed in the context of big data and the private sector, where data sources are discovered, ingested, processed, stored, analyzed, and ultimately exploited by the organization to add value. [3] The falling cost of electronic storage and the exponential growth of data have been well documented, but what has spurred the data revolution is not the volume of data but the recognition that the data have valuable uses. Finding high-value uses and creating a process to transform raw data into actionable information is the essence of the data value chain. In an increasingly data-driven world, data have been described as the new oil. [4] Ultimately, when data are put to use, they have an impact: a decision is made, a condition is altered, and someone’s well-being is affected.
In this paper, we describe a value chain for data that are usually described as official statistics, produced by governments or other public agencies. At each stage we include examples of important, value-adding activities. The examples come from desk research on the potential impact of gender data.
The production and interpretation of statistics derived from large datasets is a complex process. For the data user at any stage of the value chain, there must be confidence – trust – that the data are fit for their intended purpose. This is an example of the classic principal-agent problem, in which the agent – the data producer – has information that the principal – the user – needs to assess the value of the data. Another way to describe this problem is “asymmetric information.” Data quality assessment frameworks help to overcome information asymmetries by providing metadata about each stage of the data production process. [5] Trust in data and the interpretation of data are essential at each step of the data value chain.
Without trust, there can be no value.

