Research Support Handbook

Pseudonymisation and Anonymisation

Policies and Legislation
Research Data

If your research involves data about human beings, these data probably will be identifiable, meaning that this information can reveal your research participants’ identities. Take a look at the image below:

Image of brown chicken in a context of white chicken, explaining that even a general description can reveal someone’s identity

This example (Hrudey et al. 2019) shows that even an apparently general description can reveal someone’s identity, depending on the context. This is not problematic in itself, it means that you must make sure to carry out your research in line with the GDPR. Having said that, it is nevertheless useful to see if and to what extent you can make your data less identifiable. We explain why below.

Different levels of identifiability

Within the GDPR definitions several terms are used: pseudonymisation, anonymisation, direct identification and indirect identification. All of these terms are related to the extent to which it is possible to identify an individual. Pseudonymisation is a process to make personal data less easily linkable to individual data subjects or research participants. In other words, it is a method to de-identify personal data. If the data undergo enough de-identification that it is no longer possible to re-identify a data subject, they are considered anonymised (see also Support et al. (2025)). This process is called anonymisation.

The processes of pseudonymisation and anonymisation are depicted in the image below.

Table with fully identifiable data in the first column, pseudonymous data in the second column and anonymous data in the third column

This overview (Hrudey et al. 2019) explains the difference between fully identifiable, pseudonymous and anonymous data and provides an example of how data can be made less identifiable. The data in the left-most column are fully identifiable. The information is made less identifiable in the second column, for example by replacing the patient number with a random study subject number, and by aggregating some of the data, for example by using year of birth instead of the specific date. The combination of variables in the second column still makes it possible to reveal this person’s identity, but it is more difficult. In the third column, the pieces of information are aggregated even further, making it impossible to identify the person. These data are considered anonymous, but note that this information is probably too general for many scientific research purposes.

Why are pseudonymisation and anonymisation useful?

Full anonymisation is not always achievable or the steps involved may render the data less useful for analysis. The extent to which you will de-identify your data depends on:

  • Characteristics of the dataset
  • The context in which it was obtained
  • What the researcher plans to do with the data
  • The resources available for making the data less identifiable

Even if you cannot fully anonymise your data, a basic level of data pseudonymisation, such as removing names and contact information from a dataset, has important advantages. Pseudonymisation helps you to:

  • Safeguard the privacy of research subjects, which helps maintain public trust
  • Prevent developing a bias when working with the data
  • Meet data protection obligations
  • Decrease the privacy risks posed by your data which:
    • Increases your data storage options
    • Allows you to more securely share data with appropriate parties

Pseudonymisation and anonymisation methods

As there are many different types of data in very different formats, there is no uniform method to apply data pseudonymisation and anonymisation. General considerations for making data less identifiable are provided in the Guide ‘How can you pseudonymise and anonymise personal data?’. This Guide includes references to more specific recommendations for specific types of data (e.g. audiovisual data, consent forms, imaging data, tabular data and questionnaire data).

Acknowledgement: This text is based on the Data Privacy Handbook of Utrecht University (Support et al. 2025) and the FGB (VU Faculty of Behavioural and Movement Sciences) Security Tips. We thank our colleagues for creating and sharing their work.

References

Hrudey, E. Jessica, Jan Lucas van der Ploeg, Joan Schrijvers, et al. 2019. Anonymization - Reference Card for Researchers. Version final. Zenodo. https://doi.org/10.5281/zenodo.3584842.
Support, Research Data Management, Dorien Huijser, Neha Moopen, et al. 2025. Data Privacy Handbook. Zenodo. https://doi.org/10.5281/zenodo.15350653.