OpenAI is now providing API customers with the ability to watermark text output from selected models, a move aimed at meeting the requirements of the European Union's AI Act, which mandates that generative AI service providers make machine-readable identifiers on model output. Within weeks, the company plans to apply the watermarking automatically to ChatGPT and Codex text output within the EU, though customers outside Europe can opt to have the marks applied globally.
Anthropic was the first major AI company to announce its approach to AI Act compliance, using Google's SynthID-Text algorithm on a global basis. Microsoft and Meta have also been developing similar technology for image-based AI due to broader concerns about content provenance and misinformation.
OpenAI's watermarking method, detailed in a paper titled textGrain, works by periodically replacing a predicted word token with a synonym, creating a deviation from the unbiased model's statistical pattern. A watermark detector, available only to approved researchers and organizations, checks for that signal.
However, the technique has notable limitations. The target error rate is one percent, but in passages of 200 words, the detector flags only about 80 percent of watermarks, compared to 95 percent in 400-word passages. Performance is worse on functional texts like math or code, where synonym substitution is more likely to introduce errors. The paper also does not address whether altered word choices could change the meaning of a passage.
The watermark can be easily defeated. In tests on 400-token passages, swapping out 10 percent of words with synonyms reduced detection from 92 percent to 66 percent, and replacing 25 percent of words lowered detection to just 17 percent.
OpenAI already applies other content provenance signals to AI-generated images, audio, and video. The textGrain approach, while meeting the letter of the AI Act, may prove less effective in practice than some alternatives.