Ten principles for machine-actionable data management plans

4 min read
Curated from journals.plos.org →

Data management plans (DMPs) are documents accompanying research proposals and project outputs. DMPs are created as free-form text and describe the data and tools employed in scientific investigations. They are often seen as an administrative exercise and not as an integral part of research practice.

There is now widespread recognition that the DMP can have more thematic, machine-actionable richness with added value for all stakeholders: researchers, funders, repository managers, research administrators, data librarians, and others.

The research community is moving toward a shared goal of making DMPs machine-actionable to improve the experience for all involved by exchanging information across research tools and systems and embedding DMPs in existing workflows.

This will enable parts of the DMP to be automatically generated and shared, thus reducing administrative burdens and improving the quality of information within a DMP. This paper presents 10 principles to put machine-actionable DMPs (maDMPs) into practice and realize their benefits.

The principles contain specific actions that various stakeholders are already undertaking or should undertake in order to work together across research communities to achieve the larger aims of the principles themselves. We describe existing initiatives to highlight how much progress has already been made toward achieving the goals of maDMPs as well as a call to action for those who wish to get involved.

Data management plans (DMPs) are documents accompanying research proposals. They describe the data that are used and produced during the course of research activities, where the data will be archived, which licenses and constraints apply, and to whom credit should be given.

DMPs are awareness tools to help researchers manage their data and ensure that it will be of high quality, accessible, and reusable after the project has ended. DMPs are typically created manually, mostly by researchers using checklists and online questionnaires. They are required by funding bodies and institutions all over the world, e.g., the National Science Foundation (NSF) in the United States, the European Commission in Europe, and the National Research Foundation (NRF) in South Africa. The current manifestation of a DMP—a static document often created before a project begins—only contributes to the perception that DMPs are an annoying administrative exercise and do not support data management activities. Questions can remain unanswered, or the answers can be overly generic due to the use of free-form text.

What DMPs really are, or at least should be, is an integral part of research practice, because today most research across all disciplines involves data, code, and other digital components (often in addition to physical materials, which can also be described in a DMP).

A DMP describes digital research methods that will necessarily evolve over the course of a project; therefore, to be a useful tool for researchers and others, the content must be updated to capture the methods that are employed and the data that are produced. There is movement in this direction, e.g., Horizon2020 in Europe requires a DMP with varying levels of detail at different stages of a project, but this remains based on static text files.

We continue to need a human-readable narrative, but there is now widespread recognition that the DMP could have more thematic, machine-actionable richness with added value for all stakeholders. This includes funders, repository managers, administrators, researchers, and so on (Fig 1)—in short, everyone who is part of the larger ecosystem in which data are produced, transformed, exchanged, and reused. Stakeholders with a role in realizing the maDMP vision. Funder: funding agencies and foundations that specify requirements for DMPs and monitor compliance. Ethics review: IRBs/REBs that authorize human subjects research. Legal expert: technology transfer offices; copyright and patent lawyers.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at journals.plos.org →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.