Skip to main content
Praxikon

Guideline

Guidelines 03/2026 on web scraping in the context of generative AI

Date
Status
under consultation
Body
European Data Protection Board (EDPB)
Reference
Guidelines 03/2026, versie 1.0

What it is about

Version 1.0 was adopted on 7 July 2026 for public consultation; comments can be submitted until 30 October 2026 (23:59 CET). The guidelines cover private entities scraping data from websites to train generative AI, either themselves or through a contracted party. They address roles (controller or processor), purpose limitation, transparency (including the Article 14(5)(b) disproportionate effort exception), data minimisation, accuracy, legitimate interest as legal basis with the balancing test and mitigating measures, and special categories of data, where the GC and Others ruling (C-136/17) can be relevant for incidental collection.

What this means in practice

If you scrape web data for AI yourself on the basis of legitimate interest, you must be able to demonstrate the balancing test, and it helps to exclude sources and data types you do not need, skip websites that refuse scraping and give people a way to object in advance. If you use an already scraped dataset, you as controller assess whether it can lawfully be used and must meet the accountability principle. You must actively try to keep out special category data. It is still a draft; the final text may differ.

The GDPR articles concerned

Source: EDPB public consultation pagechecked on 15 September 2026

Summary and practical reading by Praxikon. Not legal advice; the source prevails.

Connections

What connects to this development

The counterpart in the other law

Case law8 of 15

Guidelines8 of 16

Enforcement and fines8 of 13

Legislation in motion6 of 9