OpenAI adds text watermarking for EU AI Act compliance, but technique has limitations

The company offers API customers the option to watermark AI output and will apply it to ChatGPT and Codex text in the European Union within weeks

By LineZotpaper
Published
Read Time2 min
OpenAI has begun offering API customers the option to watermark text output from certain models to comply with the European Union's AI Act, and plans to apply the watermarking to ChatGPT and Codex output in the EU in the coming weeks. The technique, called textGrain, adds an invisible statistical signal to word choices, but tests show it can be defeated by simple synonym substitution, and detection rates drop significantly on shorter passages.

OpenAI is now providing API customers with the ability to watermark text output from selected models, a move aimed at meeting the requirements of the European Union's AI Act, which mandates that generative AI service providers make machine-readable identifiers on model output. Within weeks, the company plans to apply the watermarking automatically to ChatGPT and Codex text output within the EU, though customers outside Europe can opt to have the marks applied globally.

Anthropic was the first major AI company to announce its approach to AI Act compliance, using Google's SynthID-Text algorithm on a global basis. Microsoft and Meta have also been developing similar technology for image-based AI due to broader concerns about content provenance and misinformation.

OpenAI's watermarking method, detailed in a paper titled textGrain, works by periodically replacing a predicted word token with a synonym, creating a deviation from the unbiased model's statistical pattern. A watermark detector, available only to approved researchers and organizations, checks for that signal.

However, the technique has notable limitations. The target error rate is one percent, but in passages of 200 words, the detector flags only about 80 percent of watermarks, compared to 95 percent in 400-word passages. Performance is worse on functional texts like math or code, where synonym substitution is more likely to introduce errors. The paper also does not address whether altered word choices could change the meaning of a passage.

The watermark can be easily defeated. In tests on 400-token passages, swapping out 10 percent of words with synonyms reduced detection from 92 percent to 66 percent, and replacing 25 percent of words lowered detection to just 17 percent.

OpenAI already applies other content provenance signals to AI-generated images, audio, and video. The textGrain approach, while meeting the letter of the AI Act, may prove less effective in practice than some alternatives.

§

Analysis

Why This Matters

  • The EU AI Act is the first major regulatory framework for AI, and how companies comply sets a precedent for global regulation. OpenAI's approach shows the tensions between compliance and practical effectiveness.
  • The watermark's limitations mean it may not reliably prevent misuse of AI-generated text for misinformation or fraud, particularly in short passages or specialized domains like code.
  • The contrast between OpenAI's EU-only application and Anthropic's global roll-out highlights different strategies for handling regulatory and societal expectations around AI content provenance.

Background

The European Union's AI Act, which came into force in stages starting in 2024, requires providers of general-purpose AI models to ensure that machine-generated content is identifiable. This is intended to help combat disinformation, fraud, and other harms. Major AI companies have been developing methods to meet this requirement, with OpenAI focusing on text watermarking and competitors like Anthropic adopting different algorithms. OpenAI has already applied similar provenance techniques to images, audio, and video.

Key Perspectives

OpenAI: Argues that its textGrain watermarking meets the requirements of the EU AI Act while offering customers flexibility. The company notes the approach is a work in progress and that the detector is available to approved researchers. Anthropic and other rivals: Anthropic has adopted a more universal approach, applying watermarking globally using Google's SynthID-Text algorithm. This suggests a view that content provenance is a global issue, not limited to Europe. Critics and researchers: The technique's relatively low detection rate on short passages, vulnerability to synonym substitution, and potential for unintended meaning changes raise questions about its real-world utility. The performance on code and math texts is a particular concern for developers and technical users.

What to Watch

  • Whether OpenAI extends watermarking globally in response to pressure or competitive moves from Anthropic and others.
  • The European Commission's response to the effectiveness of watermarking techniques, and whether minimum detection thresholds are set in the AI Act's implementing rules.
  • Further research into watermark robustness, including tests of more sophisticated attacks beyond synonym substitution.

Sources

Zotpaper

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.