EU AI Act and web data collection for AI training and scraping teams

The EU AI Act is now the rulebook that web-data and AI teams can’t ignore. From August 2026 it gains real enforcement teeth, and two of its obligations land directly on how training data is collected: providers of general-purpose AI (GPAI) models must publish a summary of their training data and run a copyright policy that respects machine-readable opt-outs. In plain terms, the signals a website uses to say “don’t mine me” — robots.txt, ai.txt, TDM reservations — now carry legal weight for anyone training models for the EU market. This guide explains what changed, who’s affected, and what to do if you collect web data.

I’m Andrii Byzov, an AI-Native Fractional CMO who works with web-data pipelines daily. Below is a practical, 2026-current overview — with the important caveat that this is general information, not legal advice. For adjacent reading see robots.txt and AI crawlers and is web scraping legal.


Key Facts

  • The EU AI Act (Regulation 2024/1689) is the world’s first comprehensive AI law — risk-based, phased in from 2024 to 2027.
  • GPAI rules started 2 Aug 2025; enforcement teeth arrive 2 Aug 2026 (GPAI fines up to €15M or 3% of turnover, plus high-risk and transparency obligations).
  • Training-data summary: GPAI providers must publish an overview of their training content using the Commission’s mandatory template.
  • Copyright / TDM opt-out: providers must have a policy that identifies and respects rights reservations made under the EU text-and-data-mining exception — including via machine-readable signals like robots.txt and ai.txt.
  • It targets model providers, not proxies: the obligations attach to who builds/places GPAI models on the EU market and how data is collected — infrastructure like proxies is neutral; what matters is what you collect and whether you honor opt-outs.

What the EU AI Act is — and why it touches web data

The EU AI Act is a risk-based regulation: it sorts AI uses into prohibited, high-risk, limited-risk and minimal-risk tiers and sets obligations accordingly. Most of it is about how AI systems are built and deployed — but a specific slice targets general-purpose AI models (the large models behind chatbots and assistants), and that slice reaches back into the data-collection stage. Because modern models are trained on web-scale data, the Act’s transparency and copyright rules effectively regulate how that web data is gathered and documented.

The timeline that matters

Date What applies
1 Aug 2024 AI Act enters into force
2 Feb 2025 Prohibited-practice bans + AI literacy duties
2 Aug 2025 GPAI obligations begin; training-data-summary requirement takes effect
2 Aug 2026 Enforcement teeth: GPAI fines, transparency rules, high-risk obligations
2 Aug 2027 Legacy GPAI models (placed before Aug 2025) must publish their training-data summary

Two obligations that hit data collection

1. The training-data summary. Under Article 53, GPAI providers must publish a “sufficiently detailed summary” of the content used to train the model, using a template released by the Commission’s AI Office. It asks for the types of content, the data sources, and the collection methods — with more detail expected for publicly scraped data than for licensed or private datasets. The practical effect: if you scrape the web to train a model, you need to know and document where your data came from.

2. The copyright and TDM opt-out policy. Also under Article 53, providers must maintain a policy to comply with EU copyright law and, in particular, to identify and respect rights reservations made under the text-and-data-mining (TDM) exception of the DSM Directive (Article 4). Rightsholders can reserve their works from mining, and — crucially — those reservations are expressed through machine-readable signals. The GPAI Code of Practice commits signatories to respect robots.txt and machine-readable rights-reservation protocols. That’s the bridge between a legal duty and your crawler config.

Machine-readable opt-outs now carry legal weight

This is the single most important change for data teams. For years, robots.txt was purely a norm (see our robots.txt and AI crawlers guide). Under the AI Act’s copyright regime, honoring machine-readable opt-outs — robots.txt, ai.txt, and emerging TDM-reservation protocols — becomes part of a GPAI provider’s compliance story. Ignoring a valid, machine-readable “do not mine” signal isn’t just impolite anymore; it can undercut the copyright-compliance policy the Act requires. The Commission has even consulted on standard protocols for expressing these reservations.


Who this actually applies to

  • GPAI model providers — anyone building or placing a general-purpose model on the EU market carries the training-data-summary and copyright-policy duties, regardless of where they’re headquartered.
  • Teams training or fine-tuning models for EU users on scraped data — you inherit the documentation and opt-out expectations.
  • Data suppliers and scrapers feeding those pipelines — your collection practices become part of a customer’s compliance chain, so provenance and opt-out handling matter commercially.

Note the overlap with GDPR: if scraped data contains personal information, EU data-protection law applies on top of the AI Act. The two regimes stack.

What to do if you collect web data for AI

  • Honor machine-readable opt-outs. Read and respect robots.txt and TDM/ai.txt reservations for the sources you touch — it’s now part of copyright compliance, not just etiquette.
  • Keep provenance. Record where each dataset came from, when, and how it was collected, so you can populate a training-data summary and answer downstream questions.
  • Prefer public, non-personal data and screen for personal information to manage GDPR exposure; don’t bypass logins or access-controls.
  • Collect ethically and transparently. Use consented, ethically sourced infrastructure and reasonable rates — the sourcing of your tools is part of the story you’ll be asked to tell.

Where proxies fit

Proxies are neutral infrastructure: the AI Act regulates what you collect and whether you honor opt-outs, not the transport. That said, the sourcing of your proxies is part of a defensible, ethical pipeline — which is why ethically sourced residential IPs matter. Proxies also help on the compliance-testing side: because rules, consent banners and content differ by country, checking how your own site (or a target) presents across EU markets from local IPs gives you an accurate picture. DataImpulse proxies are ethically sourced and pay-as-you-go, with country-level targeting across the EU — the clean-infrastructure layer under a responsible collection process, not a substitute for the legal work.


FAQ

Does the EU AI Act ban web scraping?

No. It doesn’t ban scraping. It regulates general-purpose AI model providers — requiring a training-data summary and a copyright policy that respects machine-readable opt-outs — which indirectly shapes how training data is collected and documented.

When does the EU AI Act start being enforced?

It entered into force on 1 August 2024 and phases in: GPAI obligations from 2 August 2025, and enforcement teeth (GPAI fines up to €15M or 3% of turnover, transparency and high-risk obligations) from 2 August 2026.

What is the training-data summary?

Under Article 53, GPAI providers must publish a sufficiently detailed overview of the content used to train their model, using the Commission’s mandatory template — covering content types, sources and collection methods, with more detail for publicly scraped data.

Do I have to respect robots.txt under the AI Act?

Effectively yes, if you train models for the EU market. The Act’s copyright rules require respecting text-and-data-mining opt-outs expressed through machine-readable signals, and the GPAI Code of Practice commits to honoring robots.txt and rights-reservation protocols.

Does the AI Act apply to non-EU companies?

Yes. The obligations attach to providers placing general-purpose AI models on the EU market regardless of where they are based, so non-EU teams serving EU users are in scope.


Build your AI data pipeline on clean, ethical infrastructure

Responsible AI data collection starts with honoring site signals and running on ethically sourced IPs. Get ethically sourced residential proxies from $1/GB — pay-as-you-go, EU and global coverage, with the geo control to collect and test responsibly.

This article is general information, not legal advice. Consult qualified counsel about your specific obligations under the EU AI Act, the DSM Directive and GDPR.

Share article: