Processing of personal data in the context of artificial intelligence models

02 July 2025

Artificial intelligence increasingly relies on personal data – but is it always in compliance with the GDPR? The European Data Protection Board (EDPB) in Opinion 28/2024 specifies when AI models can be considered anonymous, what obligations data controllers have, and what penalties may arise for violations. If you are creating or implementing AI models using personal data – this article will help you understand what conditions you must meet to operate legally and avoid substantial fines. Check how the EDPB interprets key provisions of the GDPR in the era of artificial intelligence.

The EDPB opinion was issued in response to the following questions from the Irish supervisory authority:

  1. when and how can an AI model be considered anonymous?
  2. how can data controllers demonstrate the appropriateness of legitimate interest as a legal basis during the development and implementation phases of an AI model?
  3. what are the consequences of unlawful processing of personal data during the development phase of an AI model for subsequent processing or operation of that model?

The EDPB opinion is directed to the supervisory authorities of individual EU member states (in Poland – to the President of the Polish DPA). In practice, however, it also serves as guidance for data processors involved in AI models. Therefore, this article presents the EDPB's position regarding the legality of such processing.

Which AI models are covered by the EDPB opinion

The EU Artificial Intelligence Act defines an artificial intelligence system as “a machine system designed to operate with varying levels of autonomy after its deployment and which may exhibit adaptive capabilities after its deployment, and which – for the purpose of explicit or implicit goals – infers how to generate outcomes such as predictions, content, recommendations, or decisions based on received input data that may affect the physical or virtual environment” (Article 3(1) AI Act). Thus, a key feature of AI systems is their ability to infer.

Although AI models are essential components of artificial intelligence systems, they do not constitute these systems by themselves. For an AI model to become an AI system, additional elements must be added to it, such as a user interface. AI models are typically integrated into AI systems and form part of them. The scope of the EDPB opinion covers only a subset of AI models that result from training such models using personal data.

AI models and the definition of personal data

The GDPR defines personal data as any information relating to an identified or identifiable natural person (i.e., the person to whom the data relates). To determine whether a person is identifiable, all reasonably likely means that could be used by the data controller or someone else must be taken into account.

AI models, regardless of whether they are trained using personal data or not, are typically designed to predict or draw conclusions. Moreover, AI models trained using personal data are often designed to infer conclusions about individuals other than those whose personal data was used for training. However, there are instances where some AI models are specifically designed to provide personal data concerning the individuals whose personal data was used to train the model, or to share such data in some manner.

Managing Compliance with AI

In such cases, AI models will inherently contain information relating to an identified or identifiable natural person, and therefore will involve the processing of personal data. This applies, for example, to a generative model fine-tuned based on an individual's voice recordings to mimic their voice.

Research on the extraction of training data shows that in certain cases it is possible to employ means that are highly likely to allow for the extraction of personal data from some AI models or simply to inadvertently obtain personal data as a result of interacting with the AI model.

Based on the above considerations, the EDPB indicates that AI models trained on personal data cannot in all cases be considered anonymous. The assessment of whether an AI model is anonymous should be made on a case-by-case basis, taking into account specific criteria.

When can an AI model be considered anonymous

AI models typically do not contain data that can be directly extracted or linked, but rather parameters representing probabilistic relationships between the data contained in the model. Nevertheless, in realistic scenarios, there is a risk that specific information can be inferred from the model.

Therefore, the supervisory authority, in agreeing with the data controller that a given AI model can be considered anonymous, should verify at least whether it has received sufficient evidence that, using reasonable means:

  1. all information regarding the Data Protection Impact Assessment, including any assessments and decisions in which it was determined that a Data Protection Impact Assessment was not necessary;
  2. any advice or opinions provided by the Data Protection Officer (if appointed or should have been appointed);
  3. information on the technical and organizational measures taken during the design of the AI model to reduce the likelihood of identification, including the threat model and risk assessments on which these measures are based. This information should include specific measures for each source of training data sets, including relevant source URLs and descriptions of the measures taken (or already taken by third-party data providers);
  4. technical and organizational measures taken at all stages of the AI model's lifecycle that contributed to the absence of personal data in the model or confirmed this absence;
  5. documentation demonstrating the theoretical resilience of the AI model against re-identification techniques, as well as control measures designed to limit or assess the effectiveness and impact of attacks (such as regurgitation, exfiltration, etc.). This may include, in particular:
    • the ratio of the amount of training data to the number of parameters in the model, including impact analysis,
    • re-identification probability indicators based on the current state of knowledge,
    • reports on how the model was tested (by whom, when, how, and to what extent),
    • test results;
  1. documentation provided to the data controller implementing (or implementing) the AI model or to the data subjects, in particular documentation regarding the measures taken to reduce the likelihood of identification and regarding any residual risks.

DPO Function - it is well transferred

Legitimate interest as a legal basis for data processing

The GDPR does not establish any hierarchy among the different legal bases specified in Article 6(1). To determine whether the processing of personal data can be based on Article 6(1)(f) of the GDPR, supervisory authorities should check whether data controllers have accurately assessed and documented whether the following three conditions are cumulatively met:

  1. the controller or a third party is pursuing a legitimate interest,
  2. processing is necessary for the purposes of legitimate interests,
  3. the legitimate interest is not subordinate to the interests or fundamental rights and freedoms of the data subjects.

Condition 1: existence of an interest

The interest of the data controller or a third party refers to a broader interest or benefit that they may derive from engaging in a specific processing activity. An interest can be considered legitimate if the following three criteria are met cumulatively:

  1. the interest is lawful,
  2. the interest is clearly and precisely formulated,
  3. the interest is real and present, not speculative.

Examples of legitimate interests include: the development of a virtual consultant service providing assistance to users, the development of artificial intelligence systems for detecting fraudulent content or behavior, and the improvement of threat detection in an information system.

Condition 2: determining whether the processing of personal data is necessary for the realization of the interest

The necessity test involves determining whether:

  1. the processing activity will achieve the purpose,
  2. there is no less intrusive means to achieve that purpose.

The intended amount of personal data involved in the AI model should be assessed in relation to less intrusive alternatives that may be reasonably available to achieve the purpose of the legitimate interest just as effectively. If achieving the purpose is also possible using an AI model that does not involve the processing of personal data, it should be concluded that the processing of personal data is not necessary.

Condition 3: balancing interests

The third stage of the legitimate interest assessment is "balancing" (also referred to in the opinion of the EDPB as the "legitimate interest assessment"). This stage involves identifying and describing the various opposing rights and interests that are engaged, i.e., on one side the interests, fundamental rights, and freedoms of the data subjects, and on the other side the interests of the data controller or a third party. Then, the specific circumstances of the case should be considered to demonstrate that the legitimate interest is an appropriate legal basis for the processing activities in question.

What are the interests of the data subjects

GDPR Tools
Working with good GDPR tools is not work!
Applications, calculators, GDPR snapshots - everything that can facilitate your management of the personal data protection system.
SEE MORE
The interests of individuals whose data is being processed are interests that may be affected by the processing being considered. In the context of the AI model development phase, these may include, among others, the interest in self-determination and maintaining control over one's personal data (e.g., data collected for the purpose of developing the model). Conversely, in the context of the implementation of the AI model, the interests of individuals whose data is being processed may include, among others, interests related to maintaining control over one's personal data (e.g., data processed after the model has been implemented), financial interests (e.g., when the AI model is used by the individual whose data is being processed to generate revenue or is used by a natural person in their professional activity), personal benefits (e.g., when the AI model is used to improve the accessibility of certain services), or socio-economic interests (e.g., when the AI model enables access to better healthcare or facilitates the exercise of a fundamental right, such as access to education).

Legitimate Interest and Impact Assessment of Processing

The impact of processing on individuals whose data is being processed may depend on:

  1. the nature of the data processed by the models,
  2. the context of the processing,
  3. the further consequences that the processing may have.

With regard to the nature of the processed data, it should be recalled that in addition to special categories of personal data and data concerning criminal convictions and violations of the law, which are subject to additional protection under Articles 9 and 10 of the GDPR, the processing of certain other categories of personal data may have serious consequences for individuals whose data is being processed. In this context, the processing of certain types of personal data revealing highly private information (e.g., financial data or location data) for the purpose of developing and implementing an AI model should be regarded as potentially having a serious impact on individuals whose data is being processed.

In relation to the context of processing, it is essential to first identify the elements that may pose a risk to the individuals whose data is being processed (e.g., the manner in which the model was developed, how the model may be implemented, or whether the security measures employed to protect personal data are adequate). It is also necessary to assess the significance of these threats to the individuals concerned. For instance, the use of web scraping during the AI model development phase may lead – in the absence of sufficient safeguards – to a significant impact on individuals due to the large volume of data collected, the high number of individuals whose data is involved, and the mass collection of personal data.

When assessing the impact of processing on the individuals concerned, it is also important to consider the further consequences that processing may entail. The analysis of possible further consequences of processing should also take into account the likelihood of their materialization. For example, supervisory authorities may consider whether measures have been implemented to prevent the misuse of the AI model. In the case of AI models that may be deployed for various purposes, such as generative artificial intelligence, this may include controls aimed at minimizing their use for harmful practices, such as creating deepfakes.

Legitimate Interest vs. Reasonable Expectations of the Individuals Concerned

According to Recital 47 of the GDPR, establishing the existence of a legitimate interest would require a careful assessment, including determining at the time and in the context of the collection of personal data whether the individual whose data is being processed can expect that processing for this purpose may occur. The interests and fundamental rights of the individual concerned may outweigh the interests of the data controller, particularly when personal data is processed in a situation where the individuals concerned do not expect further processing. For example, the mere fact that information regarding the AI model development phase is included in the data controller's privacy policy does not necessarily mean that the individuals concerned can reasonably expect their personal data to be used for this purpose.

Legitimate Interest vs. Risk Mitigation Measures

If the interests, rights, and freedoms of the data subjects appear to be overriding the legitimate interests pursued by the data controller or a third party, the data controller may consider implementing risk mitigation measures to limit the impact of processing on those individuals. Risk mitigation measures are safeguards that should be tailored to the circumstances of the case and depend on various factors, including the intended use of the AI model. Such measures would aim to ensure that there are no overriding interests of the data subjects over the interests of the data controller or a third party, so that the data controller can rely on this legal basis.

Examples of risk mitigation measures may include pseudonymization, measures aimed at masking personal data or replacing it with false personal data in the training dataset, adhering to a reasonable period between the collection of the training dataset and its use, and in the context of web scraping – ensuring that certain categories of data are not collected or excluding certain sources from data collection (this may include certain websites that are particularly invasive due to the sensitivity of their subject matter).

GDPR E-learning is already a standard!

Possible impact of unlawful processing during the development of the AI model on the legality of subsequent processing or the operation of the AI model

It is worth recalling that in the event of a violation, supervisory authorities may impose remedial measures, such as ordering data controllers, taking into account the circumstances of each case, to take action to rectify the illegality of the initial processing. These measures may include, for example, the imposition of an administrative fine, the imposition of a temporary restriction on processing, ordering the deletion of parts of the dataset that were processed unlawfully, or if this is not possible, depending on the factual circumstances and considering the proportionality of the measure, ordering the deletion of the entire dataset used to develop the AI model or the deletion of the AI model itself.

The EROD has outlined three scenarios illustrating the possible impact of unlawful processing during the development of the AI model on the legality of subsequent processing or the operation of the AI model.

Scenario 1:

The data controller unlawfully processes personal data for the purpose of developing an AI model, whereby the personal data is stored in the model and subsequently processed by the same data controller (e.g., in the context of implementing the model). Whether the phases of development and implementation are associated with distinct purposes (and thus constitute separate processing activities), as well as the extent to which the lack of a legal basis for the initial processing activity affects the lawfulness of subsequent processing, should be assessed individually for each case, depending on the context of the matter.

Scenario 2:

The data controller unlawfully processes personal data for the purpose of developing an AI model, whereby the personal data is stored in the model and processed by another data controller in the context of implementing the model. Determining the roles assigned to these different entities within the framework of data protection is a necessary step in identifying which obligations arising from the GDPR apply and who is responsible for them. Furthermore, in assessing the obligations of each party under the GDPR, situations of joint controllership should be taken into account. Supervisory authorities should consider whether the data controller implementing the model has conducted an appropriate assessment as part of its accountability obligations and to demonstrate compliance with the GDPR. Regarding the potential impact of the unlawfulness of the initial processing on subsequent processing conducted by another data controller, such an assessment should be carried out by supervisory authorities individually for each case.

Scenario 3:

The data controller unlawfully processes personal data for the purpose of developing an AI model, and then ensures the anonymization of the model before the same or another data controller begins further processing of personal data in the context of implementing the model. Supervisory authorities are competent and have the authority to intervene regarding processing related to the anonymization of the model, as well as processing conducted during the development phase. Consequently, supervisory authorities may, depending on the specific circumstances of the case, impose remedial measures on this initial processing.

Summary

AI models trained using personal data are fully subject to the provisions of the GDPR. Although these requirements may pose significant challenges for the development of such technologies, their primary aim is to protect the rights and freedoms of natural persons. The regulations are intended to prevent abuses related to the use of personal data during the training of artificial intelligence models. Will the mechanisms adopted prove effective in practice? The answer to this question will come with time.

Read also:

Receive a free package of 4 tutorials and 4 e-learning trainings
The controller of your data is ODO 24 sp. z o. o.