Glossary

Authors
Affiliations

Max Planck Institute for Human Development

Tobias Bengfort

RDI

Josefine Blunk

RDI

Neele Engelmann

CHM

Thomas Feg

SCT

Stefan Herzog

ARC

Maike Kleemeyer

RDI

Sina Schwarze

LIP

Sebastian Nix

RDI

Aaron Peikert

LIP

Ilse Pit

ARC

Penelope Tilsley

CEN

Published

July 12, 2026

Anonymous data
Data that can not be linked to a natural person, not even by using additional data sources. See https://gdpr-info.eu/art-4-gdpr/ for further information.
Castellum
Castellum is a turnkey open-source web application for the data protection-compliant management of participants and their personal data that has been developed at the MPIB. See here for further information.
Codebook
A codebook is a separate sidecar file that accompanies a data file. It describes the contents of the data file in detail, i.e., the variables in the data, their units, value labels, etc. and follows a predefined format.
Data Management Plan (DMP)
A data management plan serves as a central document that includes all relevant information relating to the management of research data generated in the course of a project. Among other things, it includes key content elements, such as a project description, methodology, data and associated metadata, as well as data structuring, storage, and security.
FAIR principles
FAIR is an acronym for Findable, Accessible, Interoperable, and Reusable. The FAIR data principles state that it should be possible to find research data, there should be information about how to gain access to them, they should be compatible with other data, and it should be possible to reuse them.
General Data Protection Regulation (GDPR)
The GDPR is the main piece of EU legislation governing the protection of personal data. It applies to all organizations that process data of EU citizens, regardless of where the organization is based (cf. https://gdpr-info.eu).
Hybrid journals
In addition to predominantly access-restricted articles in one journal, hybrid journals also publish individual Open Access articles that are only released in return for payment of a publication fee. All other articles can only be read in return for a fee. The publisher therefore retains the often criticized subscription models while obtaining an additional source of income through so-called “double dipping” for individual Open Access articles. In the sense of this guideline, hybrid journals are only journals which are not part of transformative agreements signed by the MPG.
Metadata
Data are usually not self-explanatory, but require additional information, so-called metadata. Metadata summarize basic information about the data, so one can immediately understand what they cover and how they have been created or collected. They are usually provided via a README and/or Codebook, as well as, e.g., on a repository’s website.
Open license
Open Licenses are a set of conditions applied to an original work that grant permission for anyone to make use of that work as long as they follow the conditions of the license. The copyright owner – usually the creator of the work, whether an individual, a group, or a company/organization – can choose to openly license their work if they want others to be able to use it freely, build on it, customize it, or improve it. There are several open licenses that follow these principles, among the most common are Creative Commons licenses. See here for more information.
Persistent identifier
Persistent identifiers allow the unique and lasting identification of a digital object or a contributor/author, e.g., DOI (digital object identifier), ORCID (Open Research and Contributor Identifier), or the Research Organization Registry (ROR).
Pseudonymous data
Data that can not be linked to a natural person directly, but can be linked by using additional data sources. Example: Data that are saved with Castellum pseudonyms instead of names (and are otherwise anonymous). See here for further information.
README-file
This is a text document that provides a clear and concise description of all relevant details about data collection, processing, analysis, and naming conventions. It should be easy to write and easy to read in order to be understandable by yourself and others in the future.
Research Data Management (RDM)
RDM refers to all measures relating to research data. This includes planning, collecting, storing, analyzing, describing, archiving, and sharing research data as well as their re-use.
Research product
The result of a scientific process is a research product. This includes, for example, but by no means exclusively text publications (e.g., articles, books, conference papers, talks, research reports, contributions to edited volumes, scientific monographs), data, and software but may also include talks, posters or websites.
Study Registration Tool (SRT)
The SRT provides a central overview of studies involving data collection conducted at the MPIB. Study registration allows the RDM Team to provide customized support throughout the data life cycle. You can register your studies here.

Reuse

Citation

BibTeX citation:
@misc{max_planck_institute_for_human_development2025,
  author = {{Max Planck Institute for Human Development} and Bengfort,
    Tobias and Blunk, Josefine and Engelmann, Neele and Feg, Thomas and
    Herzog, Stefan and Kleemeyer, Maike and Schwarze, Sina and Nix,
    Sebastian and Peikert, Aaron and Pit, Ilse and Tilsley, Penelope},
  title = {Glossary},
  version = {1},
  date = {2025-09-26},
  url = {https://os-rdm.mpib.berlin/guidelines/},
  doi = {10.17617/2.3682163},
  langid = {en}
}