Article 50
What the EU AI Act requires for marking AI text
The EU AI Act tells providers of generative systems to mark synthetic output in a machine-readable form that can be detected as artificial. It does not hand every reader a detector, and it does not turn a missing mark into proof that a person wrote the page. ChatGPT’s textGrain rollout is one provider’s answer for text.
This is an explainer for editors and developers, not legal advice. The regulation text is Regulation (EU) 2024/1689. A Commission paper on the technical options for Article 50(2) text is on the Publications Office. OpenAI’s implementation note is our approach to EU text provenance. If you need a position for your own product, that is a lawyer’s job. OpenAI says the same thing to its API customers.
What Article 50 is asking for
Article 50 is the transparency article. For generative systems, the provider-side duty that watermark discussions actually cite is the requirement to mark outputs as artificially generated or manipulated, in a machine-readable format, so they can be detected as such. A second, related duty falls on people who publish deepfakes or AI text on matters of public interest: a visible disclosure in many of those cases. The machine-readable mark and the sentence a human reads are not substitutes.
OpenAI’s help center repeats that distinction. Watermarks and C2PA give software something to read. They do not replace a banner or other notice a particular use may require. A blog that ships textGrain and no byline disclosure has not automatically finished the publisher’s problem. A blog that adds a byline disclosure and strips every file credential has not automatically finished the provider’s problem either. They are different sentences in the article.
The regulation is aimed first at providers of the systems, not at every freelancer who pastes a paragraph. Downstream duties still exist, and they depend on how the text is used. “The model was supposed to watermark, therefore my reprint is compliant” is not a reading you should rely on without advice. The practical consequence for a site like this one is narrower: do not treat a cleaner as a compliance tool.
Marking methods the Commission groups
The Publications Office report on Article 50(2) groups technical approaches into watermarking, structural marking, metadata, logging, and after-the-fact detection of AI text. It then scores them against effectiveness, robustness, reliability, accessibility, and interoperability. That list is useful because it stops the argument that there is a single compliant implementation.
- Watermarking embeds a signal in the output. textGrain and SynthID-Text are this row. They can survive light edits. They fail on short, low-entropy, or heavily rewritten text. Detection is not something every citizen can run if the key is restricted.
- Metadata is C2PA and the older EXIF/XMP tags. Rich, and easy to drop with a screenshot or a re-save. The C2PA guide is this row.
- Detection after the fact is the classifier row: GPTZero, Turnitin’s AI writing model, and the rest. The Commission paper treats this as a different methodology from an embedded mark. OpenAI does too. See watermark versus an AI detector.
- Logging keeps a provider-side record. It can be strong for the provider and useless to a third party holding only the paragraph.
Interoperability is the sore point. A textGrain detector does not read Claude’s key. A C2PA reader does not read either text key. The Act’s hope for a detectable mark is not the same as a hope that one website understands every vendor. The lab comparison is Claude, Gemini, and ChatGPT.
What the code of practice changes
Providers, including OpenAI and Anthropic, point at the EU Code of Practice on Transparency of AI-Generated Content, signed in 2026, as the working detail under the article. OpenAI uses it for two operational limits that searchers ask about constantly. The code, as OpenAI describes it, does not require watermarks on outputs shorter than 200 tokens, about 150 English words, or on code snippets. Anthropic’s watermark note is written against the same code.
Those limits are about the provider’s marking duty. They are not a safe harbor for a user who wanted a short answer unmarked so a classifier would miss it. Short model text can still look like model text. Code can still be model code. The exception recognizes that a sampling watermark has nothing to hold onto, which is also what the textGrain explainer says in technical language.
Robustness cuts the other way. A mark that disappears when you change one synonym would not answer the article. A mark that survives a full human rewrite would be claiming authorship the edit had already changed. OpenAI’s public position is the middle: common light edits should leave a detectable pattern, and substantial paraphrase or translation may not. Read a miss with that middle in mind. The essay on negatives is what a missing watermark means.
How ChatGPT and Claude answered
OpenAI’s answer for text is textGrain on eligible ChatGPT and Codex output in the EU, with API watermarking available worldwide as an opt-in rather than a default. Detector access starts with approved research and academic organizations, which OpenAI ties to the code’s expectations about responsible interpretation. The coverage table is does ChatGPT watermark text in 2026.
Anthropic’s answer is a SynthID-Text variant on covered Claude models, applied globally, plus C2PA on supported files. They say a regional on-switch was not a split they wanted to operate. Detection is similarly not a public paste box.
Neither answer makes Turnitin the enforcement mechanism. A school’s classifier is not Article 50 detection. That confusion is common enough to have its own page: can Turnitin detect ChatGPT watermarks.
What a mark does not settle
OpenAI is explicit, and the caution belongs next to the regulation rather than only next to the math. A watermark does not identify the user or the prompt. It does not say how much of a mixed document the model wrote. It does not say the text is true. It does not transfer copyright. It does not, by itself, tell a publisher whether a visible label is required for this use.
A missing mark does not settle the opposite. Unsupported paths, dates before the rollout, short text, code, factual strings, and rewrites are all unmarked or undetectable for boring reasons. So is text from a provider that chose metadata instead of a sampling watermark, or that has not finished a rollout.
Hidden characters are not a compliance mark. They are a clipboard accident. Cleaning them with the ChatGPT watermark detector is hygiene. It is not an Article 50 implementation and not an Article 50 evasion. The character layer is documented separately.
What builders and publishers can do now
If you ship an OpenAI-backed product and you want marked text, the documented step is the API text-provenance setting, per project or per organization, on the models listed there. Keep a record of when you enabled it. Do not promise customers a textGrain verdict unless you actually have detector access.
If you publish text or images you did not generate yourself, ask for disclosure and for original files, then check the carriers you can check. Invisible Unicode and C2PA are in reach. Vendor text keys are not. The editorial sequence is how to check a blog writer. The tool map is free checkers compared. Stripping a manifest so a platform will not see a credential, in a context where the mark was required, is the case to avoid. The removal page separates that from ordinary file hygiene: can you remove a text watermark.
Check listed carriers on a file you are allowed to handle →For the statute itself, use EUR-Lex rather than a blog recap. For the provider’s implementation, use the provider. This page is the map between them.
A publisher’s Tuesday
A magazine receives a feature, two generated portraits, and a note that says “drafted with assistance.” Article 50 does not hand the copy desk a single button. The practical split looks like this.
The portraits are files. If they still contain C2PA, a credentials reader can show the signed claim, and the desk can decide whether the claim stays attached in the CMS. If the only copy is a screenshot from Slack, the manifest is already gone and the pixels may or may not still carry a vendor watermark. That is a production fact, not a legal conclusion. The C2PA page is the checklist.
The feature is text. The desk does not have textGrain access or Claude’s detector. What it has is the writer’s disclosure, the sources, and the option to scan for invisible characters before typesetting. A visible credit or an editor’s note may be what readers, and some public-interest contexts, actually require. A quiet classifier does not retire that sentence. Neither does a noisy one. File the disclosure next to the piece. The intake habit is how to check a writer.
If someone on the desk strips a manifest so a partner platform will not show a badge, stop and ask whether that badge was the machine-readable mark a contract expected you to keep. Hygiene on a BOM in a pull-quote is not the same click.
A product team’s Tuesday
A team shipping a writing API on top of OpenAI has a different Tuesday. Their provider, for the base model, is OpenAI. Their own product may still be a system people look at under the same transparency conversation. The settings they can touch today are the text-provenance toggle and the model list. Turning it on marks eligible generations. It does not give them a detector to embed in the customer dashboard. Promising “we verify every completion with OpenAI’s watermark detector” is false unless the access grant exists.
They also have to decide what their proxy does after the tokens arrive. A translation step, a brand-voice rewrite, or a JSON repair pass can be the substantial edit that makes detection less reliable. If the product markets “watermarked output,” the watermark has to be in the bytes the customer receives, not in a pre-rewrite log the customer never sees. Record the pipeline. The engineering notes on the switch are in the 2026 rollout page.
Short UI strings are the trap. Button microcopy and 30-word summaries sit under the length floor OpenAI cites from the code of practice. Do not design a compliance story that depends on detecting those strings. Put the disclosure in the interface instead, where a person can read it, and keep machine-readable marks for the long generations where a sampling watermark has room to exist.
None of this is a substitute for reading Article 50 against the product you actually ship. The Commission’s grouping of methods is there so a team does not pretend that a classifier widget in the corner is the same control as an embedded mark. Pick the row you implemented, name it, and keep the other rows from leaking into the privacy policy by accident. A privacy policy that says “we watermark with C2PA” while the product only runs a classifier, or the reverse, is the interoperability failure the Commission paper is warning about, written into your own footer.
FAQ
Does the AI Act require textGrain specifically?
No. It requires machine-readable marking that can be detected. textGrain is OpenAI’s method for eligible EU text. Other methods and other providers exist.
Are short answers and code exempt from the watermark?
OpenAI says the transparency code of practice does not require a watermark under about 200 tokens or on code snippets. That does not make the text human.
Does the watermark replace a visible label?
No. OpenAI says software-readable marks do not replace banners or other notices a use may require.
Is this legal advice?
No. Read the regulation and get advice for your product and your jurisdiction.