Swap one word in ten for a synonym, and the odds of catching AI-generated text drop from 92% to 66%. That's the number OpenAI itself published while rolling out textGrain, a new invisible watermarking system for ChatGPT and Codex in the European Union.
There's no visible marker to strip out. A secret key nudges the model's word choices just slightly, hundreds of times over a passage, leaving a statistical fingerprint invisible to readers but detectable by an algorithm. OpenAI built textGrain with researchers from the University of Pennsylvania and Yale.
The timing isn't a coincidence. The EU's AI Act transparency rules, in force since August 2, require companies to label AI-generated content in a verifiable way. Over the coming weeks, ChatGPT and Codex users in the EU — on every plan, free through enterprise — will start getting watermarked output. Developers using the API can already switch it on worldwide for select models, though it stays off by default and isn't becoming a global standard yet.
The detector itself stays locked down for now: only vetted researchers and organizations get access, not any employer or teacher who wants to run a check. And the method has real limits. On 400-token passages, accuracy hits 95% at a 1% false-positive rate — but swap a quarter of the words for synonyms and that falls to 17%. Short replies, math answers and translated text are harder to flag, and a missing watermark proves nothing about human authorship.
Anthropic rolled out something similar for Claude two months earlier, and promptly got pushback from users worried about being outed at work. OpenAI seems to have seen that coming: it's going out of its way to stress the watermark reveals nothing about the account, the prompt, or the conversation behind it — just that a model wrote the words.



