OpenAI text watermark

How the OpenAI ChatGPT text watermark actually works

OpenAI’s text watermark is a pattern in which words the model picks, not a hidden character you can highlight. The system is called textGrain. This page explains the mechanism in ordinary language, using OpenAI’s own provenance notes and the textGrain technical report published with its EU text-provenance update.

Three things people call a ChatGPT watermark

Search results mash these together. They are not interchangeable, and a tool that handles one of them is silent on the others.

What people search What it actually is What a check can say
ChatGPT hidden characters, zero-width space Real Unicode codepoints that can ride along in a paste. Invisible on screen. Not OpenAI’s text watermark. A character scan can count them. See ChatGPT hidden characters.
ChatGPT text watermark, textGrain, EU AI Act mark A secret pattern in token choice, spread across the wording. Only a detector with OpenAI’s key and settings. Not a public paste box.
Turnitin, GPTZero, “AI score” A classifier that looks at writing patterns after the fact. OpenAI describes this as different from an embedded signal. The vendor’s own report. Cleaning invisible characters does not rewrite the prose those models score.

Where textGrain is turned on

OpenAI’s help center, updated alongside the October 2026 provenance notes, describes the rollout this way:

A watermark is evidence that an OpenAI model likely generated or processed the content. OpenAI says it does not identify the user, the prompt, who owns the text, or whether the text is accurate. No detection is not proof that a person wrote it.

How the next word gets chosen

A language model does not look up a single correct next word. At each step it assigns a likelihood to many possible next pieces, called tokens. Normally it samples from that list.

textGrain, as OpenAI describes it, builds several adjusted copies of that list. Each copy nudges probability toward a different set of tokens. The copies are balanced so that, averaged together, they still match the model’s original likelihoods. A secret key picks which copy to sample from. The model then makes a fresh random draw from that copy.

One sentence looks ordinary. Across a passage, the key’s preferred choices show up more often than chance. A detector that has the same key and matching settings tests that surplus. It does not need the prompt, and the technical report says it does not need to know the entropy budget used when the text was generated.

The report frames the adjustment as an optimal-transport problem: pair tokens with keyed randomness, using costs from Gumbel variables and a Kullback–Leibler penalty. The KL gap is the information the watermark spends. In plain terms, that penalty is the “knob” for how strong the pattern is. A stronger pattern is easier to detect and spends more of the model’s freedom to phrase the same idea in different ways. OpenAI says that on the benchmarks it published for Astra, and on earlier ChatGPT thumbs-down rates, the quality change sat inside ordinary noise.

OpenAI also says detection is uneven across languages. In one published chart, at a 1% false-positive rate, Spanish was the easiest of the 24 official EU languages (69.0%) and Romanian the hardest (42.2%) before they turned the knob up for languages under 60%.

What the watermark does not do

Who can actually detect it

This is the part most “ChatGPT watermark detector” pages skip. OpenAI says the text detector is not included with the API switch. Access is for approved research and academic organizations studying whether the mark works and whether people understand the result. Applications are reviewed one by one. Enabling watermarking on your own API project does not hand you the detector.

The public page at openai.com/verify checks supported image and audio files for OpenAI provenance signals such as SynthID or a trusted C2PA manifest. It is not a box where you paste an essay.

Third-party products such as Turnitin’s AI writing report, which Turnitin documents separately from its similarity score, classify prose. They do not read OpenAI’s key. A high score from one and a miss from the other can both be true.

What you can check on text you have

If the question is “does this paste contain the statistical ChatGPT watermark?”, a page that does not have OpenAI’s key should say so. This site does not have that key. Cleaning characters, or paraphrasing, is not a way to certify that a vendor detector will miss the text.

If the question is “did something invisible come along with the paste, or does this file carry a provenance manifest?”, that is a listed-carrier check. Use the ChatGPT watermark detector on text you are allowed to process. It counts invisible Unicode and, for an upload, C2PA and AI metadata. It leaves every visible word in place. It will not report textGrain, and a clean result does not mean a person wrote the draft.

Check text or a file for listed carriers →

Read the matching guide before you upload something you do not own: invisible Unicode in ChatGPT pastes or C2PA manifests in images. The guide index keeps the three mechanisms apart.

How this relates to older public research

The paper people still cite for “the” LLM watermark is Kirchenbauer and colleagues, 2023: split the vocabulary with a pseudorandom key, and gently prefer one half when sampling. That green-list idea is the public ancestor of this whole topic. textGrain is OpenAI’s own construction. OpenAI says it built it to control the tradeoff between detectability and response variety more tightly than SynthID for text or TextSeal, and that it plans to release the method as open source. Until that code is public, the help center and the technical report are the documents to trust over a blog that treats every scheme as one regex.

What the detector actually computes

OpenAI’s technical report is more specific than the help-center sketch, and it is worth knowing what it does not need. At each scored position the detector rebuilds the same tokenizer, the same block and column counts, the same context-window rule, and the same keyed pseudorandom construction the generator used. The observed token picks a block. The key and the preceding context pick a column. The cost of that pair is the evidence at that position. The report says this reconstruction does not need the coupling that was solved at generation time, does not need the model’s logits, and does not need the entropy budget the generator chose.

Those costs are combined into a sum. Under the report’s assumptions, if the text was not watermarked, the sum behaves like a gamma random variable whose shape is the number of scored positions. A watermark is declared when the sum is past the quantile that matches a chosen false-positive rate. That is why “1 percent false-positive rate” shows up next to the language chart: it is the threshold, not a promise that 1 percent of human essays will be accused in the wild. The assumptions include independence across distinct contexts. Short text has fewer positions, so the same threshold is harder to clear. Code and quotations produce positions the scheme does not really get to score.

None of that computation is available on a public form. Writing it out is so you can see why a regex, a zero-width scan, and a perplexity score are not approximate versions of the same test. They do not rebuild the key. The ChatGPT watermark detector on this site stops at listed characters and file metadata for that reason.

How strong the pattern is allowed to be

The help center’s “several adjusted copies of the likelihoods” is the picture. The report’s knob is how far those copies are allowed to move. The Kullback–Leibler gap between the watermarked draw and the original distribution is the entropy the watermark spends. Spend more, and the preferred tokens show up more reliably, which is easier to detect and, in principle, a larger change to the model’s habits. Spend less, and the text stays closer to the unwatermarked sampler, which is harder to detect. OpenAI’s public claim is that the operating point they ship does not move the benchmarks they published, or earlier thumbs-down rates, outside ordinary noise.

The language chart is the same knob in production. At a 1 percent false-positive target, Spanish detected more often than Romanian on their translated-prompt set. They then strengthened languages that sat under 60 percent. Detection is therefore not a single number you can quote for “ChatGPT” as a whole. It is a number for a language, a length, a topic, and a strength setting. A blog that prints one accuracy percentage without those four is inventing a precision the report does not offer.

Blocked transport is an implementation detail that explains why this can run at sampling time. Instead of solving a giant assignment over the whole vocabulary, the method transports blocks of tokens. Relative probabilities inside a block stay put. You do not need that sentence to use ChatGPT. You need it to ignore explainers that claim the watermark inserts a second hidden token or edits the string after the model has finished. The adjustment is inside the draw.

How to read a result if you ever get one

If an approved detector returns a hit, the help center’s limits still bind. The hit is evidence that an OpenAI model likely generated or processed the content. It is not a user id, not a prompt, not a percentage of the document, and not a ruling that the text is true. A hit on a paragraph inside a longer report leaves the rest of the report unmeasured unless those spans were scored too.

If it returns a miss, read what a missing watermark means before you announce a human author. The short version is the help center’s list: predates the rollout, unsupported path, too short, code, highly factual, or rewritten. Coverage, which is the question of whether the path was even supposed to be marked, is the 2026 rollout page. Other labs’ keys will miss on purpose. That comparison is Claude, Gemini, and ChatGPT.

If what you have today is a character count or a C2PA report, label it as that. Those checks are worth running on files you own. They are not a preview of the gamma test above. Cleaning a zero-width space does not change the sum the detector would have computed, because the sum is over tokens, not over format characters. The removal page keeps those operations apart, and the checker comparison is there so a style score does not get filed under this heading by mistake.

FAQ

Does ChatGPT add hidden characters as its text watermark?

No. OpenAI says textGrain changes word choice statistics. It does not insert hidden characters, invisible spaces, or odd punctuation. Those characters can still show up as a paste artifact. They are a separate layer.

Can I detect the OpenAI text watermark myself?

Not with a public text checker at launch. OpenAI limits the text detector to approved research and academic organizations. Image and audio checks are the ones on openai.com/verify.

Does no watermark mean a human wrote it?

No. The text may predate the rollout, come from outside the EU ChatGPT path, come from an API project that never opted in, be too short, be code, or have been rewritten. OpenAI states this limit directly.

Will removing invisible Unicode remove textGrain?

No. Deleting zero-width characters does not change the words that carry the statistical pattern, and it does not change how a style classifier scores the prose.

Does the watermark name the user or the prompt?

No. OpenAI says provenance signals do not include the user, the organization, or the prompt.

← All AI watermark guides