Research Support Handbook

How can you pseudonymise and anonymise personal data?

Process & Analyse
Learn how you can make data about humans less identifiable.

Pseudonymisation and Anonymisation

If your research involves data about human beings, these data probably will be identifiable, meaning that this information can reveal your research participants’ identities. Take a look at the image below:

Image of brown chicken in a context of white chicken, explaining that even a general description can reveal someone’s identity

This example (Hrudey et al. 2019) shows that even an apparently general description can reveal someone’s identity, depending on the context. This is not problematic in itself, it means that you must make sure to carry out your research in line with the GDPR. Having said that, it is nevertheless useful to see if and to what extent you can make your data less identifiable. We explain why below.

Different levels of identifiability

Within the GDPR definitions several terms are used: pseudonymisation, anonymisation, direct identification and indirect identification. All of these terms are related to the extent to which it is possible to identify an individual. Pseudonymisation is a process to make personal data less easily linkable to individual data subjects or research participants. In other words, it is a method to de-identify personal data. If the data undergo enough de-identification that it is no longer possible to re-identify a data subject, they are considered anonymised (see also Support et al. (2025)). This process is called anonymisation.

The processes of pseudonymisation and anonymisation are depicted in the image below.

Table with fully identifiable data in the first column, pseudonymous data in the second column and anonymous data in the third column

This overview (Hrudey et al. 2019) explains the difference between fully identifiable, pseudonymous and anonymous data and provides an example of how data can be made less identifiable. The data in the left-most column are fully identifiable. The information is made less identifiable in the second column, for example by replacing the patient number with a random study subject number, and by aggregating some of the data, for example by using year of birth instead of the specific date. The combination of variables in the second column still makes it possible to reveal this person’s identity, but it is more difficult. In the third column, the pieces of information are aggregated even further, making it impossible to identify the person. These data are considered anonymous, but note that this information is probably too general for many scientific research purposes.

Why are pseudonymisation and anonymisation useful?

Full anonymisation is not always achievable or the steps involved may render the data less useful for analysis. The extent to which you will de-identify your data depends on:

  • Characteristics of the dataset
  • The context in which it was obtained
  • What the researcher plans to do with the data
  • The resources available for making the data less identifiable

Even if you cannot fully anonymise your data, a basic level of data pseudonymisation, such as removing names and contact information from a dataset, has important advantages. Pseudonymisation helps you to:

  • Safeguard the privacy of research subjects, which helps maintain public trust
  • Prevent developing a bias when working with the data
  • Meet data protection obligations
  • Decrease the privacy risks posed by your data which:
    • Increases your data storage options
    • Allows you to more securely share data with appropriate parties

Pseudonymisation and anonymisation methods

As there are many different types of data in very different formats, there is no uniform method to apply data pseudonymisation and anonymisation. General considerations for making data less identifiable are provided in the Guide ‘How can you pseudonymise and anonymise personal data?’. This Guide includes references to more specific recommendations for specific types of data (e.g. audiovisual data, consent forms, imaging data, tabular data and questionnaire data).

Acknowledgement: This text is based on the Data Privacy Handbook of Utrecht University (Support et al. 2025) and the FGB (VU Faculty of Behavioural and Movement Sciences) Security Tips. We thank our colleagues for creating and sharing their work.

How do I pseudonymise and anonymise my data

In very general terms pseudonymisation and anonymisation involve the following steps:

  1. Write a data management plan so that you know exactly which data you need for which purposes, as well as how these data will be processed to achieve your research goals
  2. Identify any potentially directly identifying information in your data
  3. Assess whether you need to collect this directly identifying information. For example:
    1. Do you really need IP addresses in your survey data?
    2. Do you really need to record audio or video?
    3. Do you really need a consent form with a name, contact information, and signature on it?
  4. If you do not need directly identifying information to answer your research question, but you do need it to, for example, contact data subjects:
    1. Separate directly identifying information from the research data.
    2. Use pseudonyms or hashes to refer to individuals instead of names.
    3. Create a keyfile to link the pseudonyms to the names.
    4. Store the directly identifiable information and the keyfile in a separate location from the research data and/or in encrypted form.
  5. Consider which types of information may lead to indirect identification, such as demographic information (age, education, occupation, etc.), geolocation, specific dates, medical conditions, unique personal characteristics, open text responses, etc.
  6. Carry out pseudonymising and anonymising the directly and indirectly identifiable data. Methods for this are described in the the FGB De-identification Guide, particularly under step 5
  7. Go as far as you can in the de-identification process and once you’ve reached the endpoint that is feasible for your research, reassess the privacy risks posed by your data.

Pseudonymisation and anonymisation tools and software

There are also various “anonymisation” tools available online, such as OpenAire’s Amnesia for quantitative data. These tools can assist with the de-identification process and in some cases achieve anonymised data, however they do require knowledge of statistical anonymisation techniques. These tools also cannot tell you when the data are anonymous, so it can be difficult to tell if you’ve done enough to meet the GDPR’s definition of anonymised. If you wish to use such tools, it is a good idea to speak with your data steward for support.

Additional support

You can find a detailed guide on how to plan for and carry out pseudonymisation and anonymisation on the FGB De-identification Guide. This guide is focused on life sciences and social sciences data, so it may not be generaliseable to your situation.

In addition to the FGB guide, the University of Groningen has an excellent generalised overview on de-identification.

Making qualitative data, such as interview data, audio and video recordings, unstructured observations, less identifiable is often challenging, because it can lead to substantial information loss (Verburg et al. 2026). In the guide Making Qualitative Data Reusable, provided by DANS (Data Archiving and Networked Services), you can find some tools that support pseudonymisation and anonymisation of qualitative data (p. 19). The Community of Practice for Open Naturally Occurring Data has an FAQ about anonymising and pseudonymising naturally occurring data (NOD) (see the section ‘Anonymising and pseudonymising NOD’).

Lastly, it’s also a good idea to discuss your pseudonymisation and anonymisation plans with your your data steward and Privacy Champion, especially before making any assumptions that the data are anonymous.

References

Hrudey, E. Jessica, Jan Lucas van der Ploeg, Joan Schrijvers, et al. 2019. Anonymization - Reference Card for Researchers. Version final. Zenodo. https://doi.org/10.5281/zenodo.3584842.
Support, Research Data Management, Dorien Huijser, Neha Moopen, et al. 2025. Data Privacy Handbook. Zenodo. https://doi.org/10.5281/zenodo.15350653.
Verburg, Maaike, Ricarda Braukmann, and Widia Mahabier. 2026. Making Qualitative Data Reusable - a Short Guidebook for Researchers and Data Stewards Working with Qualitative Data. Version 3.0. Zenodo. https://doi.org/10.5281/zenodo.8319060.