A press publisher or holder of documentary collections who wants to know whether their site was used to train a model has, in principle, a tool: the public summary of training content that the AI Act requires from general-purpose AI model providers. The European Commission's template includes a section for the main scraped domain names. Of 24 summaries accessible as of 25 September 2026, 21 include this section and none name a site. The answers describe categories of sources, types of content or extensions such as .com. The three other documents, published by DeepSeek, do not include the section.

The obligation has applied since 2 August 2025 to models newly placed on the European market; those already marketed before that date have until 2 August 2027 to comply, and the Commission's sanction powers entered into force on 2 August 2026. The list of domains is, according to the Commission, the information that should enable rights holders to verify the use of their content themselves, without having to contact each provider. The 24 September posting of the inclusionAI/AI-Transparency repository on Hugging Face provides eleven new documents to compare, for Ant Group's Ling, Ring and Ming families, and follows the Commission's template. The corpus examined also includes two summaries from OpenAI, one from Mistral AI, one from Z.AI, four from ByteDance, two from MiniMax and three from DeepSeek. This comparison does not constitute a census of all available summaries.

The template asks for identifiable domains

Article 53, paragraph 1, point (d), of the AI Act requires general-purpose AI model providers to publish a sufficiently detailed summary of their training content. The Commission's template distinguishes public datasets, data obtained from third parties and content collected directly from the internet by the provider or on its behalf.

This last category is covered by section 2.3. The template asks for a list of first- and second-level domains, for example example.com, corresponding to the top 10% of all collected domains, ranked by volume of extracted content. An adapted threshold is provided for SMEs. The list may be provided in a downloadable file or via a link.

The explanatory notice refers to a "descriptive and synthesised form" to take account of trade secrets. However, in its note 16, it recalls the requirement for a list of domain names. The Commission's FAQ also presents this list among the information intended to help rights holders identify the origin of training content.

Source categories in the 21 sections

Ant's eleven documents use the same introductory response. The main collected domains are described as a set of web platforms, code hosting services, academic repositories, encyclopaedic platforms and regional portals. The following paragraphs specify the sectors, languages and geographical areas covered. Seven responses repeat the same text; the other four adapt the types of content mentioned. None provides a site name in this section.

The finding is repeated in the summaries for GPT-6 Astra and GPT-5.5. OpenAI lists the same categories: academic, technical, legal or government resources, document sharing services, community sites and regional portals. The response in the MiniMax M3 summary repeats this wording. That of H3 merely states that the sources cover a wide range of domains and content.

The Mistral Large 3 summary mentions generalist sites, specialised resources and government, administrative or legal domains. In that of the GLM-5 family, Z.AI adds extensions, including .com, .org and .net, to categories such as code hosting and scientific resources. An extension alone does not identify a site.

The four ByteDance summaries examined cover Seed 2.0 Pro, Seedance 2.0, Seedance 2.5 and Seedream 5.0 Pro. The responses mention either sectors, such as education and research, or the subjects represented in the content, such as objects, people and natural scenes. Again, no identifiable domain in the field provided for this purpose.

DeepSeek does not include the online collection section

The summaries for DeepSeek-V3.1, V3.2 and V4 have another particularity. After data obtained from third parties, they go directly to user data. The section devoted to web scraping and its field for domain names are absent.

These documents nevertheless describe, in the part relating to rights reservations, a collection robot designed to respect robots.txt instructions and other applicable protocols, without bypassing captchas, paywalls or password protections. They therefore provide information on the declared collection rules, without providing the list of domains concerned.

What the summaries still provide

These documents provide other usable information. Ant's document on Ling-2.0 names Ant Spider, specifies a collection period and cites Common Crawl among the public datasets. DeepSeek's summaries also cite Stack Exchange. The absence noted here concerns the section on directly collected domains, and not any reference to a source throughout the documents.

For a publisher, the consequence is precise. A category such as general knowledge resources does not allow them to find their own domain, nor to conclude that it was used or that the collection was unlawful. The explanatory notice provides a complementary mechanism: providers are invited to respond, on a voluntary basis, to rights holders who wish to know whether content from their domains has been collected. It is this recourse to the provider that the list was intended to make unnecessary in the first instance. Whether these sections satisfy Article 53 is for the Commission to decide, whose oversight powers over general-purpose models have been applicable since 2 August.

Our articles will then appear first in Google Top Stories.