Implementation

Authors
Affiliations

Max Planck Institute for Human Development

Tobias Bengfort

RDI

Josefine Blunk

RDI

Neele Engelmann

CHM

Thomas Feg

SCT

Stefan Herzog

ARC

Maike Kleemeyer

RDI

Sina Schwarze

LIP

Sebastian Nix

RDI

Aaron Peikert

LIP

Ilse Pit

ARC

Penelope Tilsley

CEN

Published

July 12, 2026

1 Managing research data

Managing research data comprises general principles but also data type specifics. The general principles that apply to all research data are included in the first section. The second section contains concrete measures for all data types relevant to the Institute. If they cannot be followed for good reason, it is recommended to consult the Institute’s RDM team and individual adaptations should be properly documented in the README file to enable reuse. The term “must” in this section is used to signify a technical requirement to enable automation of workflows or legal requirements that remain binding regardless of the guideline nature of this document.

1.1 Data Management Plan

A data management plan (DMP) is one single document, containing all information regarding the data management in a research project. It describes how data are handled during and after the end of a project. Creating a DMP is helpful for improved knowledge management (information is structured, collected centrally, and can easily be shared), compliance with good scientific practice and easier subsequent use of research data. At the very least, it is recommended that a DMP includes sections on the handling of personal data, safe and secure storage as well as considerations on data publication. In addition, considering the whole data lifecycle will help researchers to anticipate important aspects before things go wrong (e.g., you store data on a device that does not provide enough space, or you do not get consent for publishing data). Therefore researchers are required to provide a DMP, and store it as a document called “DMP” in the project folder. Usually, researchers have already done this as part of the (Internal) Review Board (i.e., ethics) application, so you can just copy it from there. Research centers may provide model data management plans for simple projects. Larger projects are recommended to create and regularly update a data management plan by using the MPG’s RDMO platform. In that case please add the RDM coordinator to their RDMO project as a member. Data that are part of larger cooperation projects or consortia may obviously share the same DMP or even refer to other institutions regarding data handling.

1.2 Handling personal data

Given the Institute’s research focus, many data that we acquire are of a personal nature. This comes with the obligation of handling the data in compliance with strict data protection laws (GDPR). To facilitate these processes, the Institute has developed and implemented a central software tool, Castellum. All researchers are therefore required to enter the individuals that have participated in their studies in Castellum if contact data are available (i.e., participant name, address, email, phone number) and not to store any contact information elsewhere. Castellum generates study-specific pseudonyms, that are used to store the actual research data. This ensures that (directly identifiable) contact information is strictly separated from research data and they can only be connected via Castellum.

Moreover, GDPR-compliant handling of research data is ensured by the following measures:

  • Suitable technical and organizational measures to protect personal data (e.g., access rights, encryption, password protection).
  • The legal basis for managing personal research data within the project is the so-called broad informed consent (cf. GDPR recital 33).
  • A data processing agreement is concluded when passing on personal data to service providers (e.g., labs, transcription services).
  • Research data can either be published in anonymized form or with the consent of the affected parties. In other words, identifiable data can not be published without consent.
  • Even, the storage (without publication) of identifiable data requires good reason, therefore pseudonymization and maximimum storage of ten years for such data is recommended.
  • The legal basis as well as any necessary measures to be taken are agreed with the Institute’s Data Protection Coordinator.
  • Partiticipant consent forms will be made available to everyone involved in the handling of identified data.

1.3 Storage

All research data and materials relevant to the study and reproducing its results are recommended to be stored on servers hosted by MPG preferably or GWDG. Highly sensitive data like brain data, genetic data, and patience data must be stored on servers hosted by MPG or GWDG. Ideally, the data is stored on MPIB servers in the study folder that is created via the study registration tool. It is recommended that a subfolder is created either directly as ???-study/data/ or in the private folder as ???-study/private/data/ if personal data are contained. If data are stored elsewhere, the location should be referenced in the study registration tool.

Using the institutional infrastructure will ensure suitable technical and organizational measures to protect personal data (e.g., access rights, encryption, password protection) as well as regular automated back-ups.

1.4 Access management

All Institute members have read-access to all study folders and subfolders except the private folder. Hence, if your project includes personal data, these must be stored in the private folder to be GDPR-compliant. Access to the study’s private folder is managed via the study registration tool and can be updated whenever needed.

1.5 Documenting data

Documenting data enables adequate reuse by providing information on the context and the content of the data. The data are documented in one or more separate files, for which suitable templates are available. It is recommended that the context of the data is specified via a README file. A README template is automatically placed in every study folder. As a minimal requirement, this README template should be completed. In order to understand the contents of data files, a “codebook” is often required, describing the contents in details, i.e., the variables in the data, their units, value labels, etc. To keep data and documentation structured, such “codebooks” are recommended to be stored as separate files, so-called sidecar files, but keeping the name of the corresponding data file, e.g.,

survey_data.csv -> data file,
survey_data.yml or survey_data.json -> codebook.

1.6 Archiving

1.6.1 Digital data

In favor of sustainably handling storage space and at the same time adhering to the MPG’s Rules of Conduct, research data will be stored for a defined period of time (usually ten years) at the MPIB’s digital archive, once the study ends. If the study is registered in the study registration tool, this process is automatically triggered upon the study’s end date. The contact person will be notified about the archival in advance. In all other cases, researchers are expected to approach the RDM Team when a study ends (i.e., when analyses are completed, manuscripts are accepted, and a period of seven days to data access can be tolerated) to trigger this process manually and free up storage capacity. The MPIB’s digital archive refers to a digital archive that takes over those digital data that no longer need to be maintained on costly, in-house storage. Digitally archived data can be restored within about a week if needed.

1.6.2 Analog data

Although most of the data we acquire is intrinsically digital, some data may also exist in analog form, such as filled out paper-based questionnaires. In order to adhere to the MPG’s Rules of Conduct, the original analog data are stored for a defined period of time (usually ten years) at the MPIB’s paper archive. In case this is needed, researchers are expected to trigger this process via the RDM team, to ensure it is well documented and connected to the study that the data were acquired for. The MPIB’s paper archive is usually carried out via an external service provider that takes over those analog data for secure archival. Please also see the section on managing analog data.

1.7 Deletion

At the end of the defined retention period, research data will be deleted from the MPIB archive. The contact person will be notified about the deletion in advance.

2 Data type specific regulations

This section contains concrete measures for all data types relevant to the Institute.

2.1 Default

If your data do not fit into any of the more specific categories below, the default regulation detailed in this section applies.

Definitions:

  • One or more investigations form a coherent research project.
  • Each investigation produces one data set — hereafter referred to as a “set” — which is the combination of all variables and observations analyzed together in the research product (e.g., demographic and task variables). Variables or observations that have not been analyzed together may form a set but are not required to (e.g., different experiments that have used the same sample but whose variables have not been analyzed together or different subpopulations that have been analyzed separately).
  • A variable is a single measured attribute or characteristic that can vary across observations (e.g., age).
  • An observation is a single case or unit of analysis in the dataset (e.g., a participant or a trial).
  • The data/ folder must only contain set-related files.

2.1.1 Investigation

  • Each investigation must have a unique name within the project (e.g., “study1” or “strooptask”).
  • Only alphanumeric characters and _ (underscore) are permissible; a - (dash) must not be used.
  • The name must be used as a prefix in all filenames for files related to the set including but not limited to the data/ folder (e.g., study1_trialdata.tsv).

2.1.2 Set

  • For each set three files must exist:
    • readable data file
    • tidy data file
    • sidecar file
  • Each set can include:
    • source data file
    • public data file

2.1.2.1 Source Data (optional)

The original software output (often in a proprietary format) that may require conversion before analysis. For example, a raw XML or JSON file directly exported from a survey or experiment software or a binary file (e.g..psydat from PsychoPy) or audio, video etc. files encoded in proprietary formats. If there is more than one file, combine them into a single zip file. If data already comes in non-proprietary formats, then it does not need to be deposited as source data, but can be deposited directly as readable data (see below).

Naming:

  • prefix:
  • suffix: “-source”
  • file extension: any
  • regex: ^<investigation_name>-source\..+$
  • example: study1_trialdata-source.psydat

2.1.2.2 Readable Data (refers to raw data in BIDS)

This file contains the minimally processed data, that is recommended to be preserved in an open-source representation such as .csv, .tsv, .mp4 or .json. No variable selection or processing apart from conversion to an open-source format is permissible. Readable data are always considered identified data and may contain personally identifiable information and may therefore not be version controlled (otherwise GDPR cannot be complied with). If there is more than one file, combine them into a single zip file.

Naming:

  • prefix:
  • suffix: “-readable”
  • file extension: any
  • regex: ^<investigation_name>-readable\..+$
  • example: study1_trialdata-readable.csv

2.1.2.3 Tidy Data

The tidy data together with the sidecar file (see below) are at the heart of the default data regulation and form the basis of all analyses conducted on this data. The tidy data file therefore contains in all conscience a final, clean, read-only version of the data (e.g., correct IDs, correct number of trials per person). Providing scripts (and other documentation) for any changes that have been performed to produce tidy data from source and/ or readable data is highly encouraged. Tidy data follow the tidy data principles (Wickham, H. 2014. Tidy Data. Journal of Statistical Software, 59(10), 1–23.):

  • Each variable is a column.
  • Each observation is a row.
  • Each cell contains an atomic value.

Such a format is often called a “long” format (although there are datasets, such as logs, that become shorter if converted to tidy data).

Do not use:

id day1 day2 day3
A  1    3    2
B  6    5    4
C  6    4    2

Instead, use:

id day happiness
A  1   1
A  2   3
A  3   2
B  1   6
B  2   5
B  3   4
C  1   6
C  2   4
C  3   2

Each set corresponds to exactly one tidy data file (one tidy data file ↔︎ one set ↔︎ one investigation). Sometimes this requires repetition of values (in the example above, the values in the id column are repeated) which is acceptable even if this increases file size and thus storage demands. The file must be a .tsv (i.e., a CSV with separator \t to avoid the common confusion between different locales (German vs US) delimiters: ,, ., and ;). More technically in terms of the CSVW/ Data Package Table Dialect, the file must conform to:

{
  "header": true,
  "headerRows": "1",
  "headerJoin": " ",
  "commentChar": "#",
  "delimiter": "\t",
  "lineTerminator": "\r\n",
  "quoteChar": "\"",
  "doubleQuote": true,
  "skipInitialSpace": false,
}

Furthermore:

  • Text encoding must be UTF-8
  • Missing data must be coded as NA.
  • Numeric values must be valid numbers in terms of JSON, i.e., conform to the regular expression ^-?[0-9]+(\.[0-9]+)?([eE][+-]?[0-9]+)?$, which most importantly entails a . as the decimal sign.

Naming:

  • prefix: <investigation_name>
  • suffix: none
  • file extension: .tsv
  • regex: ^<investigation_name>\.tsv$
  • example: study1_trialdata.tsv

2.1.2.4 Sidecar file

Each tidy data file must be accompanied by a sidecar file describing its content in detail. Each column of the tidy data file requires exactly one entry in the sidecar file. A sidecar file is a human- and machine-readable codebook for the tidy data (and thus does not have to document all of the source or readable data). Each entry in the sidecar file requires at least the fields: title (a human readable version of the variable name), and type (any of string, number, integer, boolean, date, datetime), and optionally description or any other fields allowed in the DataPackage Table Schema.

id:
  title: "A string uniquely identifying every participant."
  type: string
day:
  title: "Day of assessment of the participant, counted from the first day they appeared in the lab."
  type: integer
happiness:
  title: "How happy are you today?"
  type: number
  description: "We came up with the ingenious idea of asking people how happy they are. See Me, my friends, et al. (2025, 2024, 2022)"

Naming:

  • prefix: <investigation_name>
  • suffix: none
  • file extension: .yml
  • regex: ^<investigation_name>\.yml$
  • example: study1_trialdata.yml

The specification above is a simplification and strict subset of the DataPackage Table Schema. There are several competing metadata standards for tabular text data, among them Frictionless DataPackage, CSVW, and CSV Schema. The latter two have not seen widespread adoption or recent development but are backed by the World Wide Web Consortium/Internet Assigned Numbers Authority. The subset was chosen so that tooling for any of these can be provided rather simply. For example, the DataPackage json can automatically generate the below example:

{
  "profile": "tabular-data-resource",
  "name": "happiness",
  "path": "data/happiness.tsv",
  "format": "tsv",
  "mediatype": "text/tab-separated-values",
  "encoding": "utf-8",
  "hash": "sha256:7f8917a40c7176eab2f513006d4790c2dd69a6088b7c40b5c74a555eee2b4405",
  "bytes": 71,
  "schema": {
    "fields": [
      {
        "name": "id",
        "title": "A string uniquely identifying every participant.",
        "type": "string"
      },
      {
        "name": "day",
        "title": "Day of assessment of the participant, counted from the first day they appeared in the lab.",
        "type": "integer"
      },
      {
        "name": "happiness",
        "title": "How happy are you today?",
        "type": "number",
        "description": "We came up with the ingenious idea of asking people how happy they are. See Me, my friends, et al. (2025, 2024, 2022)"
      }
    ],
    "missingValues": ["NA"]
  },
  "dialect": {
    "header": true,
    "headerRows": "1",
    "headerJoin": " ",
    "commentChar": "#",
    "delimiter": "\t",
    "lineTerminator": "\r\n",
    "quoteChar": "\"",
    "doubleQuote": true,
    "skipInitialSpace": false,
  },
  "licenses": [
    {
      "name": "CC0",
      "title": "Creative Commons CC0",
      "path": "https://creativecommons.org/publicdomain/zero/1.0/"
    }
  ],
  "created": "2025-02-11T09:44:50Z"
}

2.1.3 Publication

To avoid publishing sensitive data, the following .gitignore file is recommended to be placed in the data/ directory:

# Ignore everything
*

# But not public files
!*-public.tsv
!*-public.yml
  1. If data size allowed the use of Git, the best way to publish the data is by mirroring the repository to Zenodo in order to assign a persistent identifier to the data.

  2. If data are stored on smb shares, you first want to choose an adequate repository. A good first hint may be provided by re3data.org or by simply talking to colleagues in your field to find out what they are using.

Further important steps include (cf. Open Data):

  • Choosing a license for your published work (usually CC0) so that others know how they can reuse your work.
  • Assigning a persistent identifier.
  • Completing all relevant metadata (links to paper, Git repo, preregistration, etc.).
  • Including a data availability statement in your paper.

The RDM Team is currently working on providing a comprehensive example.

2.2 Analog data

Most of the data we acquire are intrinsically digital. However, some data may also exist in analog form, such as filled out paper-based questionnaires. In such cases, digitizing these documents is highly desirable to avoid data loss. That is, it is recommended to scan the documents and save them in the corresponding project folder. In order to provide the data in reusable format, they should be (and are usually) exported to analyzable digital formats as described in Default.

2.3 Multi-task behavioral data

Multi-task behavioral data are stored in BIDS format.

2.3.1 Study setup

When researchers are programming behavioral experiments, they need to make sure that the output is stored following BIDS requirements, i.e., each participant’s data is stored separately in:

data
    sub-<label>/
        [ses-<label>/]
            beh/
                sub-<label>[_ses-<label>]_task-<label>[_acq-<label>][_run-<index>]_beh.json
                sub-<label>[_ses-<label>]_task-<label>[_acq-<label>][_run-<index>]_beh.tsv

2.3.2 Piloting

During piloting, researchers are responsible for ensuring that behavioral data as well as all additional data acquired during the session (e.g., behavioral event files, eyetracking, physiological data) are saved in the correct format. We recommend saving data locally during acquisition to minimize risks of network interruptions etc., but transferring them to the study folder (that is created upon study registration in the Study Registration Tool (SRT)) immediately afterwards to ensure they are safe and secure. To facilitate transfer without data loss rsync is recommended to be used rather than pure drag and drop. Contact the RDM Team in case you need help setting this up. Finally, the piloting phase is meant to be used to check and ensure data in the BIDS folder are complete and accurate (count the number of trials, check IDs etc.).

2.3.3 Data acquisition

During data acquisition, checks for completeness of data have to be conducted on a regular basis. It is recommended to correct human errors such as interrupted tasks, incorrect IDs, missingness without documentation, etc. During acquisition a printed protocol sheet (ask the RDM Team for an example) is recommended to be used to detail any acquisition issues, re-runs, etc. To this end, we recommend digitizing the session protocols immediately after completing the session, and integrating them in the participants.tsv file within the BIDS folder, e.g., with two extra columns per task:

ID Task1 Comments_task1 Task2 Comments_task2
sub-4z7cyl include exclude computer crashed
sub-8pgcr3 include include

2.3.4 Raw data

Within a reasonable time after conclusion of data collections, researchers are recommended to have a final, clean, read-only data folder. This folder is the rawdata that forms the basis of all analyses conducted on this data. Data files have sustainable file formats (.csv, .tsv) and corresponding sidecars.

Any changes that have been performed on the data derived from the behavioral test session to arrive at this final version have to be documented, e.g., in the BIDS changes file with scripts to perform the changes collected in a corresponding Git repository. If the data does not contain personally identifiable information (otherwise GDPR cannot be complied with) and size of the data set permits, data management is recommended to be performed in Git allowing for version control of the dataset. However, Git is currently not compatible with our smb shares. Thus, when using Git, your local repository should be in Seafile to still have it on our institutional infrastructure, and the link to the repository should be entered in the study registration tool. Alternatively, data management steps can be performed in datalad allowing for version control of larger datasets in a manner comparable to Git.

2.3.5 Derived data

As the name suggests, these are versions of the data that are derived from the raw data after certain processing steps (i.e., preprocessing, or analyses) have been performed. Derived data is recommended to be saved in the corresponding derivatives folder, e.g., Project/private/data/derivatives/excludeoutliers. It is recommended to manage all preprocessing and analysis code using Git.

2.3.6 Publication

Keep in mind that data can only be published if there are no legal concerns (e.g., data are anonymized, no data license issues). If legal restrictions prohibit publication of the data, an effort should be made to create a publishable version, e.g., creating aggregated versions, removing variables or synthesizing data.

  1. In case data were managed using Git, the best way to publish these data is by mirroring the repository to Zenodo in order to assign a persistent identifier to the dataset.

  2. If data are stored on smb shares, you should first choose an adequate repository. A good first hint may be provided by re3data.org or by simply talking to colleagues from your field to find out what they are using.

Further important steps include (cf. Open Data):

  • Choosing a license for your published work (usually CC0) so that others know how they can reuse your work.
  • Assigning a persistent identifier.
  • Completing all relevant metadata (links to paper, Git repo, preregistration etc.).
  • Including a data availability statement in your paper.

2.4 EEG data

The RDM Team and EEG lab Team are currently working on refining this section and your input is highly appreciated.

EEG data is recommended to be stored in BIDS format. What this means in terms of workflows and changes needs to be determined with your help!

2.5 MR data

MR data are usually of sensitive nature so that high levels of data protection apply (see also https://open-brain-consent.readthedocs.io/en/stable/gdpr/index.html)!

2.5.1 Organization

MR data are stored in BIDS format.

2.5.2 Study setup

To significantly facilitate the conversion of data into BIDS, scanner sequences should be named ReproIn compatible, see also our documentation.

2.5.3 Piloting

During piloting, researchers are responsible for ensuring that the automatic conversion of MR data works correctly. In this pipeline, the acquired MR data are automatically converted from DICOM to NIfTI file format and structured in BIDS-compatible fashion. Data are saved in the study folder that is created when a study is registered in the SRT. All additional data acquired during the MR session (e.g., behavioral event files, eyetracking, physiological data) need to be transferred manually. We recommend scripting this processes during piloting to facilitate transfer without data loss during the data acquisition stage. Example scripts for this process (e.g., to derive BIDS-compatible event files from different presentation software) will be made available in the near future. Contact the RDM Team to seek advice. These data have to be integrated into the automatically created BIDS folder on the correct level, e.g., event files within the functional folder of each participant. This process should also be scripted. Finally, it is recommended to use the piloting phase for checking and ensuring that data in the BIDS folder are complete and accurate.

2.5.4 Data acquisition

During data acquisition, checks for completeness of data have to be conducted on a regular basis. It is desirable to correct human errors such as interrupted scans, incorrect IDs, missingness without documentation, etc.

During acquisition a printed protocol sheet (ask the RDM Team for an example) is recommended to be used to detail any acquisition issues, observed movement during MRI scans, re-runs etc. We then recommend digitizing the session protocols immediately after the completion of the session, and integrating them in the participants.tsv file within the BIDS folder, e.g., with 2 columns per sequence to detail whether (1) this sequence should be included/excluded, and (2) any comments:

ID Sequence1 Sequence1_Comments Sequence2 Sequence2_Comments
sub-4z7cyl include exclude movement
sub-8pgcr3 include re-run, trigger problem include

These columns are recommended to be updated in subsequent data quality checks (e.g., if mriqc reports show movement that was not seen during acquisition).

2.5.5 Raw data

Within a reasonable time after the conclusion of data collections, researchers should have a final, clean, read-only data folder. This folder is the raw data that forms the basis of all analyses conducted on these data. Any changes that have been performed on the data derived from the MR session to arrive at this final version have to be documented, e.g., in the BIDS changes file with scripts to perform the changes collected in a corresponding Git repository. Git is currently not compatible with our smb shares. Thus, when using Git, your local repository should be in Seafile for it to remain on institutional infrastructure, and the link to the repository should be entered in the study registration tool. Additionally, data management can be performed in datalad, allowing for version control of the dataset.

2.5.6 Derived data

As the name suggests, these are versions of the data that are derived from the raw data after certain processing steps (i.e., preprocessing, or analyses) have been performed. Derived data should be saved in the corresponding derivatives folder, e.g., Project/private/data/derivatives/fmriprep. It is recommended to manage all preprocessing and analysis code using Git. Additionally, derived data can also be managed via datalad. Given the size of MR data, many of these analysis steps are likely conducted on a high-performance cluster such as the tardis. Note that tardis does not have any backup, it is therefore crucial that after completion of an analysis, data is transferred back to the derivatives folder within the study folder.

2.5.7 Data publication

Data privacy is of particular concern for MRI data, particularly anatomical data. Firstly structural images contain facial features which can easily be used to identify the participant, e.g., through facial reconstruction. One first approach would be to skull strip data before sharing, which is easily achieved with softwares such as Freesurfer and FSL. However, it may be the case for certain datasets, e.g., those integrating MEG or EEG, or for certain analysis pipelines, that skull-stripped data would render the data unusable. As a second choice, many software exist for defacing anatomical data: pydeface, mri_deface, quickshear, deepdefacer, mridefacer. One of the most effective tools is pydeface (see Theyers et al., 2021), a python-based tool that can deface an anatomical image and apply the calculated defacing mask to other structural images in the dataset.

pydeface path/to/T1w.nii.gz --cost normmi --applyto path/to/T2w.nii.gz

Secondly, MRI data in participant space whether it be structural or functional data, can also pose risks of being identifiable. One recent recommendation is to share data in standard spaces rather than participant space wherever possible. Further details on considerations when sharing MRI data can be found on Open Brain Consent and Open WIN. Ultimately, the platforms through which you would like to share each have their own specific criteria for sharing and anonymization or pseudo-anonymization of MRI data.

Any sharing of data should also always be checked against the participant consent forms to make sure you have the right to share any particular type of data.

As a quality-check step, defaced data is recommended to be verified for successful removal of key areas used in identification such as the nose (cf. https://raamana.github.io/visualqc/gallery_defacing.html). Non-defaced data can be kept in sourcedata/ or raw/ folders that should not be shared.

A further consideration in sharing MRI data is the presence of sensitive information in JSON sidecars that accompany NIfTI files. For example, JSON sidecar files may contain sensitive information within the following fields: AcquisitionTime, InstitutionAddress, InstitutionName, InstitutionalDepartmentName, ProcedureStepDescription, ProtocolName, PulseSequenceDetails, SeriesDescription and global (cf. https://peerherholz.github.io/BIDSonym/outputs.html#metadata-tsv-files). Before sharing MR data, all JSON sidecars should be screened for such sensitive information, which should be removed.

Due to the size of imaging data, the choice of repositories is somewhat limited but you can still search re3data.org. You may want to look at the following options first: Ebrains, Gin, as well as Neuovault for brain maps.

Further important steps include (cf. Open Data):

  • Choosing a license for your published work (usually CC0) so that others know how they can reuse your work.
  • Assigning a persistent identifier.
  • Completing all relevant metadata (links to paper, Git repo, preregistration, etc.).
  • Including a data availability statement in your paper.

2.6 NIRS data

The RDM Team and NIRS data user will work on refining this section and your input is highly appreciated.

NIRS data are stored in BIDS format. What this means in terms of workflows and changes needs to be determined with your help!

2.7 Simulation data

Simulation data are artificially generated data that imitates real-world processes or systems, created using statistical or computational models.

2.7.1 Simulation Code

All source code used to generate simulated data is recommended to be published to ensure transparency and reproducibility. The previous guidelines with respect to Open Software/Code apply, especially the implementation guidelines regarding versioning, publication, reproducibility, and random number generation.

2.7.2 Data Publication

If possible, simulated data should be published according to the principles outlined above and is recommended to also follow the guidelines on Open Data. Especially if it is computationally expensive or otherwise technically difficult to re-simulate the data from the published simulation code, additional publishing of the simulated data is crucial. If limitations such as limited storage preclude publishing the full data, aggregated data is recommended to be published along with the code used for aggregation.

Reuse

Citation

BibTeX citation:
@misc{max_planck_institute_for_human_development2025,
  author = {{Max Planck Institute for Human Development} and Bengfort,
    Tobias and Blunk, Josefine and Engelmann, Neele and Feg, Thomas and
    Herzog, Stefan and Kleemeyer, Maike and Schwarze, Sina and Nix,
    Sebastian and Peikert, Aaron and Pit, Ilse and Tilsley, Penelope},
  title = {Implementation},
  version = {1},
  date = {2025-09-26},
  url = {https://os-rdm.mpib.berlin/guidelines/},
  doi = {10.17617/2.3682163},
  langid = {en}
}