EDPB Guidelines on Generative AI and Web Scraping: Anonymisation and the Handling of Publicly Available Personal Data
An explanation of the EDPB's draft guidelines on web scraping for generative AI and on anonymisation, covering their legal nature and the issues they raise for companies' data handling.
Primary sources
The announcements and documents this analysis covers.
On July 8, 2026, the European Data Protection Board (EDPB) announced in a news release that at its most recent plenary it had adopted "Guidelines on web scraping in the context of generative AI" and "Guidelines on anonymisation." At the same plenary, the final version of the guidelines on blockchain technologies, for which public consultation had already been completed, was also adopted, but this article focuses on the two documents on web scraping and anonymisation, which bear directly on the collection of training data for generative AI and on personal data protection practice. Both sets of guidelines are the EDPB's views as a supervisory body clarifying the interpretation of the GDPR (EU General Data Protection Regulation), and their legal nature differs from that of the GDPR itself and of each member state's national laws. As of the time of writing, neither document is a final version; both are at the draft stage, submitted to the public consultation procedure.
What Was Adopted (Breakdown and Legal Nature of the Documents)
Of the documents adopted in this announcement, the Guidelines on web scraping in the context of generative AI (Guidelines 03/2026) and the Guidelines on anonymisation (Guidelines 02/2026) were both adopted as version 1.0. The version history section of each guidelines document expressly states "adoption of the guidelines for public consultation," and they are open for public consultation until October 30, 2026. In contrast, the guidelines on blockchain technologies were adopted as a final version, settled after a public consultation that had already been conducted. The phrase "adopts final version" in the title of the EDPB's announcement refers to these blockchain guidelines, and it would be a mistake to read it as meaning that the final versions of the web scraping and anonymisation guidelines have also been settled; I think care is needed to distinguish these stages. The cover of each guidelines document gives an adoption date of July 7, 2026, while the EDPB's news release is dated July 8.
Both sets of guidelines are documents adopted by the EDPB under Article 70(1) of the GDPR, and they set out the supervisory authorities' interpretation of how the provisions of the GDPR apply; they are understood not to create penalties or new obligations for companies on their own. The web scraping guidelines build on EDPB Opinion 28/2024 (opinion on certain data protection aspects related to the processing of personal data in the context of AI models), adopted on December 17, 2024, and add more concrete analysis of the criteria for purpose limitation and legitimate interest. The anonymisation guidelines are positioned as updating the three anonymisation criteria set out in Article 29 Working Party Opinion 05/2014 of 2014, in light of subsequent developments in CJEU case law.
The Web Scraping Guidelines
These guidelines cover cases in which an organization carries out web scraping itself, or has another company carry it out on its behalf, for the purpose of training generative AI, and they deal only with scraping by private individuals and private companies. The GDPR applies where scraping involves the processing of personal data, such as collection, storage, organization and retrieval. Where the entity carrying out the scraping and the entity developing an AI model using the collected dataset are different, whether they are to be characterized as controller, joint controllers or processor is to be determined case by case based on the facts, such as the instruction relationship between them.
Compliance with the principles of purpose limitation and transparency is also addressed. Where individual notification of data subjects is impossible or would involve a disproportionate effort, individual notification may not be required under Article 14(5)(b) of the GDPR, but even in that case, the guidelines state that information must be made publicly available through a privacy policy or the like, including the sources of collection and how data subjects can exercise their rights. With regard to the data minimisation principle, examples given include the use of synthetic data and clarification of collection criteria before collection, and measures such as anonymisation and pseudonymisation after collection. With regard to the accuracy principle, the guidelines recommend scraping from reliable sources, recording timestamps at the time of collection, and verifying data before using it for AI training.
As for the legal basis, the guidelines state that for scraping by private companies for the purpose of training generative AI, legitimate interest under Article 6(1)(f) of the GDPR is often relied on, and that the three requirements must then be satisfied: the existence of a legitimate interest, the necessity of the processing, and a balancing test against the rights and interests of data subjects. In the balancing test, emphasis is placed on whether data subjects could reasonably have expected the scraping, and collection from sites that indicate an intention to refuse scraping through robots.txt, CAPTCHAs or the like is said to weigh against such reasonable expectation. For special categories of personal data, which correspond to sensitive personal information, processing is in principle prohibited by Article 9 of the GDPR, and in addition to an exception under Article 9(2), a legal basis under Article 6 is also required. Referring to the framework of the CJEU judgment on search engines, GC & Others (C-136/17), the guidelines state that the reasoning of that judgment may extend to scraping only where the collection is unintended and incidental and the entity takes technical and organizational measures to prevent dissemination within the scope of its responsibilities, powers and capabilities, but they emphasize that this is not a general exemption from the requirements of Article 9.
The Anonymisation Guidelines
The anonymisation guidelines set out a two-step question as the core test: whether the data relates to a natural person, and whether that natural person is identified or identifiable. If the answer to either question is negative, the data is treated as anonymous data, but the guidelines state that the conclusion of this assessment may differ depending on the perspective of each entity that may access the data, in other words, depending on which entity is used as the reference point for assessing anonymity. This point is organized in light of the ruling in Case C-413/23 P EDPS v SRB (CJEU, judgment of September 4, 2025).
As frameworks for the assessment, two approaches are presented: a "contextual approach" that takes into account differences in capability among the entities that may be involved in identification, and a "simplified approach" that assesses uniformly without taking those differences into account. The contextual approach reflects the details of the legal standard, while the simplified approach may lead to conclusions more conservative than the legal standard but enables simpler and more reliable assessments. Under either approach, if all three criteria of non-singling-out, non-linkability and non-inferability are met, the data can safely be treated as anonymous data; if any one of them is not met, further analysis is required. This is positioned as maintaining the approach of the three criteria set out in Article 29 Working Party Opinion 05/2014 of 2014 while making them more concrete in light of technological trends and CJEU case law. The guidelines state that data already determined to be anonymous under Opinion 05/2014 before the publication of these guidelines is not immediately required to be reassessed, but periodic review of re-identification risk is described as good practice.
Points of Contact for Japanese Companies (Extraterritorial Application of the GDPR and Differences from the APPI)
The first thing to check is whether there are situations in which the company falls within the extraterritorial application of the GDPR. Article 3(2) of the GDPR provides that even businesses without an establishment in the EU may be subject to it where they offer goods or services to data subjects in the EU or monitor the behavior of data subjects taking place within the EU. Where a company scrapes information including personal data from websites in the EU as training data for generative AI, or provides AI services to customers and users in the EU and analyzes their behavior, I think it is necessary to check whether this is a situation in which the GDPR requirements set out in these guidelines apply directly to the company. As explained in What Companies Should Check First on AI Services and Personal Data Protection, the starting point is to take stock of whether the company's data collection flows related to generative AI may include EU personal data.
On that basis, the differences from Japan's Act on the Protection of Personal Information (APPI) should be organized. The APPI provides that the principal's consent is not required to acquire special care-required personal information where it has already been made public by the principal, a national government body, a local government, an academic research institution or the like (Article 20, paragraph 2, item 7 of the APPI), which is an express exemption concerning the acquisition of public data. On the other hand, as the EDPB's web scraping guidelines state, there is no general exemption under Article 9 of the GDPR for the processing of special categories of personal data, and the requirements of Article 9 cannot be treated as satisfied solely because the data subject has made the data public. In addition, for ordinary personal information as well, Japanese law imposes an obligation, where personal information is acquired from someone other than the principal, to promptly notify the principal of or publicly announce the purpose of use after acquisition, except in cases such as where the purpose of use is deemed clear (Article 21, paragraph 4, item 4 of the APPI and others), and the structure of this requirement differs from the standard of impossibility or disproportionate effort in Article 14(5)(b) of the GDPR. I think it is a matter to check that a practice that has relied on the exception for special care-required personal information already made public cannot be invoked as it is when dealing with data subjects in the EU.
Items to check going forward include: taking stock of whether the company or its contractors acquire data for training or tuning generative AI from external sources, and whether such data may include information about data subjects in the EU; asking vendors that provide training data about the legal basis they rely on, measures to prevent the inclusion of special care-required personal information, and whether EU data is included; and checking whether the recipients of the company's services include customers and users in the EU. For designing the items to check with vendors, see Checkpoints for AI Vendor Due Diligence and Contract Review, and for reviewing personal data clauses in outsourcing agreements, see Review Points for Personal Data Handling Clauses. Note also that the special exception, introduced by the amendment to Japan's APPI enacted on July 10, 2026, that dispenses with consent for AI development and similar activities that can be characterized as the preparation of statistics is a domestic institutional reform concerning the regulation of third-party provision under Japanese law, and care is needed not to confuse it with the GDPR analysis set out in these EDPB guidelines, which is a separate regime. Points to check regarding the copyright aspects of training data are covered in Generative AI, Copyright and Training Data Checkpoints.
Developments to Watch
As of the date of writing, the following points remain unsettled and require ongoing monitoring. The final versions of both the web scraping guidelines and the anonymisation guidelines are expected to be prepared in light of the results of the public consultation ending on October 30, 2026, and no timing for their finalization has been indicated at present. Following the public consultation, the statements concerning the criteria for legitimate interest and the operation of the three anonymisation criteria may be revised. Further developments in the CJEU case law on which the anonymisation guidelines rely, and the status of revisions to EDPB Opinion 28/2024 and EDPB Guidelines 1/2024 (guidelines on legitimate interest), which the web scraping guidelines refer to, are also subjects to monitor.