On 5 October 2026, OpenAI published a 20-page technical report on textGrain, the statistical watermark it will apply to eligible text produced by ChatGPT and Codex in the European Union. In a blog post titled "Our approach to EU text provenance rules", the company says this marking will roll out "over the coming weeks" to comply with the AI Act. Customers of its API have been able to turn it on since 5 October, worldwide, for select models, but it will remain off by default there.
The report is co-signed by five OpenAI researchers and four academics from the University of Pennsylvania and Yale. The blog post, published at 15:00 UTC according to the company's RSS feed, states that the report will be updated "in the coming weeks". OpenAI relayed the announcement the same day on its X account.
A signal carried by word choice
A language model chooses each word by sampling from a probability distribution. textGrain replaces part of that randomness with pseudo-random values computed from a secret key and the preceding words. The report describes a partition of the vocabulary into blocks, followed by an optimal transport problem with Gumbel costs that steers the choice of block. Within the selected block, words keep their original relative probabilities.
The report is built around an "entropy budget", the share of randomness one accepts to give up in order to create the signal. The authors present it as an answer to a limitation of Gumbel-max watermarking. With the same key and the same prompt, that method can produce identical responses, which is a problem for applications that request several variants of a text. The detector needs only the text and the key. The report specifies that the theoretical false positive rate rests on independence assumptions and that a deployed key requires empirical calibration.
The signal adds no hidden characters or invisible spaces: it is carried by the words themselves. OpenAI's help center states that textGrain is designed to withstand light edits and copying and pasting. Substantial paraphrasing or translation, however, can make it undetectable.
95% at 400 tokens, 17% after a rewrite
OpenAI's blog post publishes rates measured on responses to questions from the English-language ELI5 dataset. They depend on text length and topic, with a detector tuned to target a 1% false positive rate.
| Scenario tested | Reported detection |
|---|---|
| 400-token passage on a psychology topic | about 95% |
| 200-token passage | about 80% |
| Another set of 400-token passages, 10% of words replaced with synonyms | from about 92% to 66% |
| Same set, a quarter of words replaced | 17% |
| Mathematical content | "substantially lower" in the text; about 60% at 400 tokens according to The Decoder's reading of OpenAI's chart |
The two synonym rows come from a separate evaluation, run on another set of 400-token passages, with a baseline rate of about 92%.
For other languages, the help center presents an evaluation based on 500 prompts written in English and then translated into the 23 other official languages of the Union. At the same 1% false positive threshold, detection ranges from 42.2% in Romanian to 69.0% in Spanish. OpenAI says it strengthened the signal for languages below 60%, using an adjustable strength parameter.
The 200-token threshold matches the one in the European Code of Practice on transparency of AI-generated content. The code exempts "very short text", defined as text shorter than 200 tokens, from watermarking. Above that length, marking is still required even though its reliability is lower.
Article 50: an obligation in force since 2 August
Article 50(2) of Regulation (EU) 2024/1689 covers providers of AI systems that generate text, image, audio or video. Their outputs must be "marked in a machine-readable format and detectable as artificially generated or manipulated". The obligation has applied since 2 August 2026, a deadline covered in a previous ActuIA article on the AI Act countdown.
Regulation (EU) 2026/1744 of 8 July 2026, published in the Official Journal on 24 July and known as the digital omnibus on AI, added a paragraph 4 to Article 111. According to the text published by the Commission's AI Act Service Desk, systems placed on the market before 2 August 2026 have until 2 December 2026 to comply with Article 50(2). For ChatGPT and Codex, already on the market by that date, the "coming weeks" announced by OpenAI fall within this window.
On 10 June, the Commission published the Code of Practice, which it and the AI Board assessed as adequate, followed on 20 July by guidelines on Article 50. OpenAI is among the signatories of section 1, which covers providers, in the list published by the Commission, which counted around 190 organizations at the end of July. This marking obligation comes on top of a separate one, the training data summaries required of general-purpose models.
Off by default in the API: who carries the obligation?
In the API, the watermark will remain off by default, OpenAI writes. Its help center describes how to turn it on: a "Text provenance" setting in the organization settings, under Data controls, or in a project's settings, with a choice of the models concerned. Coverage is due to extend to older models over the coming weeks. It is not a request parameter, and the API reference mentioned none on 6 October. OpenAI says it is working with cloud partners, which it does not name in its blog post, to offer the watermark on their services; The Decoder cites Microsoft Azure.
For a French company that integrates an OpenAI model into an assistant or a writing tool offered under its own name, the Commission's guidelines provide the framework. They define as a provider the organization that places an AI system on the market or puts it into service under its own name or trademark, including a chatbot developed in-house for its own use. They recall that Article 50 does not explicitly apply to general-purpose AI models. A provider may rely on the marking solution implemented upstream by the model provider, to the extent that this marking is compliant and without prejudice to its own responsibility.
Under this reading of the guidelines, which do not explicitly address the case of a disabled option, an integrator that leaves the default setting in place cannot rely on OpenAI's watermark. It must then turn it on, for a project or for the whole organization, or mark outputs by other means. Article 50 provides an exception for assistive functions for standard editing and for tools that do not substantially alter the input data. Its paragraph 4 adds a separate obligation for deployers that publish generated text to inform the public on matters of public interest, unless the text has undergone human review and is subject to editorial responsibility. The same paragraph covers deepfakes, for which Spain planned to penalize the failure to label in a draft law of March 2025.
A restricted detector, within the limits of the code
On 5 October, OpenAI opened applications for access to its detector. The blog post links this choice to the measured rates: "These limitations contribute to our decision to provide initial detector access only to approved researchers and expert organizations". The help center cites partners at Cornell, ETH Zurich and the Slovak institute KInIT. Detector access thus remains subject to application, including for API customers who turn on the watermark.
The code allows this restriction for free-form text, considered less reliable, but sets limits on it. Signatories must provide controlled, free access to market surveillance authorities and other regulators, law enforcement, the media and fact-checkers. The list also includes trusted flaggers, independent researchers, educational and research institutions and civil society. The restriction must be limited in time. The guidelines also ask providers to make the means of detection available to persons exposed to the content. A newsroom or a university is among the categories the code requires signatories to serve; an HR department wanting to check a cover letter falls into none of them.
Anthropic, Google, vLLM: other choices of scope
Anthropic announced on 14 August a watermark derived from SynthID-Text, applied globally "because we don't yet have a durable way to scope it by region". At the time, the post said Anthropic would "soon be offering a watermark detection API", according to its version archived the same day. In its current version, as on the Claude Fable 5.1 launch page published in September, Anthropic presents this API as a private preview. It is open to EU regulators, law enforcement, media, fact-checkers, independent researchers, educational bodies and civil society groups, as well as to companies required to verify the watermark for their own compliance. Google had open-sourced SynthID Text in 2024. On 24 September, the vLLM project described the integration of a Gumbel-max watermark into its inference engine, in a post signed by engineers from Mistral and Red Hat. OpenAI plans to release textGrain as open source, with no date given, and states that it matched or exceeded SynthID for text in its evaluations.
OpenAI stands apart by limiting marking to the territory where the law requires it, whereas Anthropic applies it everywhere. The two companies agree on one point: "The absence of a detected watermark does not prove human authorship," OpenAI writes. The text may be too short, edited or translated, come from an unsupported model or another vendor, or predate the marking.
