Over the coming weeks, OpenAI will begin adding an invisible watermark to text generated by ChatGPT and Codex for eligible users within the European Union, in a move intended to comply with the transparency rules set out in the European AI Act. The rollout will cover all plans, while text watermarking will not become a global default setting at launch.
The transparency rules took effect on August 2 and require AI companies to identify automatically generated content in a way that enables other systems to recognize it. Meanwhile, developers using OpenAI's API anywhere in the world can activate the feature on specific models starting now, but it is disabled by default.
How does the watermark work?
The watermark does not appear as a visible symbol or alert. Instead, it is created by subtly modifying the word choices suggested by the model. This modification leaves a pattern that a specialized detection tool can identify, and the pattern remains in the text when it is copied and pasted. OpenAI says the watermark does not reveal the user's identity, and that enabling it did not cause a meaningful change in model performance during the company's tests.
OpenAI published a technical report on its method at the same time. The method, called textGrain, was prepared by researchers from the company and the University of Pennsylvania and Yale University. According to the report, the method uses a secret key to arrange the probabilities of the next words while completing sentences, then aggregates hundreds of these modifications so that the detector can identify the pattern using the text and the key.
Detection limits after editing
The watermark does not provide conclusive evidence of the text's source in every case. In a test conducted by OpenAI, replacing 10% of the words with synonyms reduced the detection rate from approximately 92% to 66%. The company also found that short passages, mathematical answers, and translated texts were more difficult to detect.
OpenAI warns that the absence of a watermark does not prove that the text was written by a human; the text may be too short, may have undergone extensive editing, or may have been generated by an AI system from another company. The presence of a watermark may also indicate that an OpenAI system generated or processed part of the passage, but it does not determine how much human judgment, editing, or creativity went into producing it.
Why does this news matter?
The rollout gives European institutions an additional technical means of examining ChatGPT and Codex content, but it does not by itself solve the problem of proving the text's source. The susceptibility to editing and translation, as well as text length, means that detection results require careful interpretation. This led OpenAI to initially limit access to detection tools to accredited researchers and specialized organizations to assess their reliability and responsible uses.
The move follows Anthropic's announcement two months ago that it would watermark text generated by Claude worldwide, a decision that faced objections from some users who believed that their contribution to the instructions, context, and decisions made Claude a tool rather than an author. OpenAI had previously developed a technology for watermarking text but refrained from releasing it, amid concerns that users would switch to competitors that did not use watermarking, according to a report published by The Wall Street Journal in 2024.