ChatGPT text with an invisible provenance watermark represented by a subtle digital signal
Image Credit: OpenAI

OpenAI Adds Invisible Watermarks to ChatGPT Text in the EU for AI Provenance

OpenAI's textGrain system will add an invisible statistical signal to eligible ChatGPT and Codex text in the EU.

TL;DR: OpenAI will add invisible textGrain watermarks to eligible ChatGPT and Codex text in the EU, while warning that a watermark can signal OpenAI involvement without proving authorship, ownership, or accuracy.

OpenAI text watermarking is coming to eligible ChatGPT and Codex output in the European Union over the coming weeks. OpenAI says the move responds to the EU AI Act requirement for generated text to be identifiable in a machine-readable way.

The system is called textGrain. Instead of adding visible labels, hidden characters, unusual spaces, or special punctuation, it changes the statistical pattern of the model's word choices. A detector can then look for that pattern in a passage.

What OpenAI's text watermark actually does

When a model generates text, it has several possible words or word pieces to choose from at each step. textGrain subtly adjusts those choices according to a secret pattern. Across a long enough passage, the choices can form a signal that a matching detector can recognize.

That makes the system different from a normal AI-text classifier. A classifier tries to infer whether text looks AI-generated after the fact. OpenAI's watermark is embedded during generation, so detection is based on a provenance signal created by the model itself.

OpenAI says the watermark does not add hidden characters or watermark-only tokens, and ordinary readers will not see anything different in the text. Its provenance documentation describes textGrain as an invisible change to the randomness used in word selection. OpenAI's provenance guidance explains the mechanism and its limits.

The EU rollout is deliberately limited

OpenAI says eligible ChatGPT and Codex text generated in the EU will receive the watermark over the coming weeks, across all ChatGPT plans. It is not making text watermarking a global default at launch.

API customers are getting a different option. Customers globally can opt in to text watermarking for select models, with the feature off by default. OpenAI also says it is working with cloud partners to extend watermarking to eligible outputs accessed through their services.

AreaOpenAI's current approach
ChatGPT and Codex in EUWatermarking rolling out
API customersGlobal opt-in for select models
API defaultOff
Detector accessApproved researchers and expert organizations
Public text detectorNot available at launch
TechnologytextGrain

OpenAI's own testing shows the watermark has limits

The most important part of the announcement may be what happens when the text is edited.

At a target false-positive rate of 1%, OpenAI says its detector identified watermarks in about 80% of 200-token psychology passages and about 95% of 400-token passages. Detection was substantially weaker for mathematics, where there is less freedom to change wording.

Editing can weaken the signal quickly. In OpenAI's test of 400-token passages, replacing 10% of the words with synonyms reduced detection from about 92% to 66%. Replacing 25% reduced detection to 17%.

A watermark is not proof of AI authorship

OpenAI says a detected watermark can indicate that an OpenAI system generated or processed part of a passage. It does not establish who wrote the text, how much a human contributed, who owns it, or whether the text is accurate.

A missing watermark does not prove human writing

OpenAI is also explicit about the other side of the detection problem. If a detector finds no watermark, that does not prove a person wrote the passage.

Short answers may not contain enough text for reliable detection. Code and highly constrained factual answers can also give the model too little freedom to alter word choices. Translation, substantial paraphrasing, and other major edits can further weaken the signal.

That distinction matters when comparing provenance signals with AI-writing detectors. A provenance watermark answers a narrower question: whether a supported OpenAI signal can be found in the text. It does not provide a complete history of how the passage was written.

Saganote's Claude text watermark explained coverage examines the same provenance problem from Anthropic's side, while Google's SynthID watermark approach shows another model-maker's approach to identifying generated content.

OpenAI says watermarking does not reduce model quality

OpenAI reports that its tests on Astra showed no meaningful performance difference between watermarked and unwatermarked text across the benchmarks it uses to evaluate the model. The company also says text watermarking has a negligible impact on model speed.

BenchmarkUnwatermarkedWatermarked
Artificial Analysis Intelligence Index49.5749.76
AutomationBench34.09%34.86%
DeepSWE v1.172.80%71.68%
Terminal-Bench 4.053.90%56.06%
BrowseComp87.92%87.35%
GPQA Diamond94.44%93.94%

Text joins OpenAI's broader provenance system

OpenAI already uses provenance signals for other media. Supported images can carry Content Credentials and SynthID, while supported generated audio uses SynthID. The new text system extends that layered approach to another content type.

The broader OpenAI product line matters too. ChatGPT Images 2.5 adds another major generated-content workflow, while GPT-Live shows why provenance increasingly has to cover more than plain text.

OpenAI is also expanding ChatGPT into business workflows. Its Data Agent turns questions into Power BI and Tableau dashboards, showing how generated and transformed information is becoming part of practical work.

The company's wider platform strategy includes advertising as well. Saganote's coverage of ChatGPT Ads tracks that expansion, although advertising and provenance are separate systems.

Provenance is not the same as authorship

OpenAI's most useful warning is that provenance signals should be treated as evidence about a model's involvement, not as a complete authorship record.

A watermark cannot tell a reviewer whether a person wrote an outline and an AI polished it, whether an AI drafted a passage that a person heavily edited, or whether the model generated the whole passage. It also does not identify the user, account, prompt, or conversation associated with the text.

That limitation will matter in schools, workplaces, publishing, and research, where the important question is often not simply whether AI was involved, but how AI was used.

What happens next

OpenAI plans to continue testing how textGrain performs under editing and translation and says it intends to make the technology open source. Detector access will initially remain with approved researchers and expert organizations so OpenAI can study reliability and responsible use before broader access.

For now, the EU rollout is best understood as a provenance signal rather than a universal AI detector. OpenAI is adding a machine-readable marker to eligible generated text, but its own results show that text length, subject matter, and editing can all affect detection.

Frequently Asked Questions

What is OpenAI text watermarking?
It is OpenAI's method of embedding an invisible statistical signal into generated text so a detector can look for evidence that an OpenAI model produced or processed the passage.
Will every ChatGPT response have a watermark?
OpenAI says eligible ChatGPT and Codex text generated in the EU will receive watermarking as the rollout expands. The company is not making it a global default at launch.
Can the watermark prove that AI wrote an entire document?
No. OpenAI says the watermark does not measure human contribution or establish authorship, ownership, responsibility, or accuracy.
Can editing remove the watermark?
Editing can weaken detection. OpenAI's tests found substantial drops in detection after replacing portions of a passage with synonyms.
Can people see the watermark?
No. textGrain changes the statistical pattern of word choices rather than adding visible markers or hidden characters.

OpenAI's official EU text provenance announcement provides the technical results and rollout details. OpenAI's provenance guidance explains how textGrain fits with its image and audio provenance systems.


Share this
Waqas Ahmad

About Author

Waqas Ahmad

Waqas Ahmad is a technology writer at Saganote and a software professional with more than a decade of experience in the software industry. He writes about AI, software development, developer tools, and emerging technologies. With hands-on experience building and working with software, Waqas focuses on making complex technologies easier to understand and explaining how they work and what they mean in practice.