Your AI Now Hides a Watermark in What It Writes
I heard a rumor this week and did the obvious thing. I asked the AI itself.
“Do you embed hidden markers in what you write for me?”
It said yes. Then it told me three more things. It could not turn the marking off. It could not show me the mark. And it could not confirm the mark was actually there, because it has no more ability to inspect its own output than I do.
I assumed I knew where this was going. Another warning label — the technology industry’s version of the cancer notice on every parking garage in California, stamped on everything until it means nothing. I was wrong, and what I found instead is stranger and harder to defend.
Quick summary: Anthropic now hides a statistical watermark in text written by Claude, and the rest of the industry is following under a new European law. The engineering is careful — it marks text in proportion to how much the machine actually wrote, so editing your own work leaves almost nothing behind, and it will not falsely accuse you the way today’s AI detectors do. The problem is who can read it. Only the company that made the mark can check for it. You cannot, your school cannot, your editor cannot. And because no test can ever return “a human wrote this,” anyone accused has no way to prove otherwise.
What actually got turned on
Anthropic has begun marking text produced by Claude. Models launched on or after August 2, 2026 carry it from day one, and older models are being brought along behind them. There is no opt-out — no setting, no subscription tier, no developer switch.
This is not one company being unusual. Article 50 of the European Union’s AI Act took effect on that same August date, requiring AI-generated content to be machine-readable as such. Roughly 190 organizations signed a code of practice ahead of it, including OpenAI, Google, Meta, Microsoft, and Mistral. Penalties reach €15 million or three percent of global revenue. Google has quietly done this in Gemini since 2024.
China got there first, and went further. Its labeling rules took effect September 1, 2025, and require not only a hidden machine-readable mark but a visible one — a label an actual human being can see, on AI-generated text, images, audio, and video.
How it works, explained with a coin
One honest warning before I start. What follows is the concept, not the machinery. The real system involves cryptographic keys, small tournaments run between candidate words, and a statistical test to decide whether a pattern exceeds chance. I am giving you the idea behind it, deliberately simplified, because the idea is genuinely understandable and the machinery genuinely is not. The research is linked at the bottom if you want the real thing.
An AI writes one word at a time. At most points in a sentence, several words would work equally well. “The car was fast” — or quick, or speedy, or rapid. Nothing is riding on the choice. The sentence is fine either way.
Now imagine a secret rulebook that only the AI company holds. At each of those little forks, the rulebook silently sorts the equally-good options into two piles. Call them heads and tails. The AI is then nudged — gently, not forced — toward the heads pile slightly more often than chance.
Look at any single word and you learn nothing. A coin comes up heads half the time on its own. But flip it a thousand times and count: if it lands heads six hundred times, that is not luck. Somebody weighted the coin.
That is the whole trick. The watermark is not hidden in any word. There is nothing buried inside “fast” to find, search for, or delete. The watermark is a lean across thousands of small choices — a tendency, visible only in bulk, and only to someone holding the rulebook.
One more detail, because it explains a great deal: the rulebook is rewritten at every position based on the few words immediately before it. Which words count as heads in one spot has nothing to do with the next spot. That is why you cannot learn the pattern by staring at it, and why changing one word disturbs its neighbors too.
Everything confusing about this follows from the coin
Short text cannot carry it. Three coin flips prove nothing about a coin. A sentence or a text message is far too small to hold a measurable pattern. The mark needs hundreds of words before it means anything.
Code is essentially exempt. A watermark needs choices to hide in, and code mostly has none — there is one correct way to write the line and every alternative is broken. No fork in the road, no coin to flip. Anthropic says plainly that where an exact output is required, the watermark is not applied.
Hard facts carry less of it. A sentence stating a date, a price, or a version number has very little wiggle room, so the marking thins out.
Editing dilutes it, in proportion — and then some. Every word you replace is one the AI no longer chose. And because each word’s rule depends on the few words before it, changing one word also scrambles the reading of the several that follow. One edit damages roughly five words of evidence. Change a little and the pattern barely moves; rewrite most of it and the pattern is gone.
Proofreading your own writing marks almost nothing. This is the part nearly every article got wrong, mine included before I read the documentation. If you wrote the sentence and the AI fixed a typo, the AI made one choice out of forty. There was never a weighted coin to begin with. In Anthropic’s own words, “nearly all the words are the person’s, there’s very little (if anything) for the watermark to attach to.”
The mark measures how much the machine wrote. Not whether it was in the room.
The good news, which is real
You do not have the rulebook. Neither does any other human being. So when you write, your word choices land on heads and tails at random, averaging out to nothing. You cannot accidentally write in a pattern keyed to a secret you have never seen. Human writing sits at plain chance.
Compare that to the AI detectors schools and publishers have been buying since 2023. Those never had a rulebook. They guess, by measuring whether writing seems too smooth, too even, too predictable — which means they punish clarity. A Stanford study found those detectors misclassified more than half of essays by writers who learned English as a second language, while scoring nearly perfectly on native-born eighth graders.
A watermark detector does not care how you write. On the specific problem of accusing innocent people, this new approach is dramatically better than what is already installed in schools. I want to give it that credit before I take it away.
And now the part nobody mentions
You cannot read it.
The rulebook is a secret held by the company. Checking for the mark requires the rulebook — without it there is no way to know which words should have been favored, so there is no test to run. That is not a temporary gap while the tooling catches up. It is the design. If the rulebook were public, anyone could strip the mark or forge it into someone else’s writing, and the whole thing collapses.
So only Anthropic can detect Anthropic’s watermark. Only Google can detect Google’s. There will never be a browser extension, a consumer app, or an independent service that reads these — and anyone selling you one is selling the same guesswork that has been misflagging students for two years.
It gets worse when more than one tool touches a document, which is now the normal case. Say an AI drafts something, you rewrite a third of it, a grammar tool cleans it up, and a phone assistant rephrases a paragraph. That document may now carry three faint marks from three companies, each weakened by the others, each readable only by its own maker. No party sees more than their own fragment. There is no registry, no shared standard, and no way to know which companies to even ask.
The contrast with images makes this hard to excuse. Pictures and files got C2PA — an open standard, signed metadata, verifiable by anyone with public tooling, one format across every vendor. Text got proprietary silos. Same regulation, same deadline, two opposite architectures, and the one covering the most content is the one nobody outside the companies can read.
Accused, with nothing to show
Here is where this stops being a technical curiosity.
Think about what a person actually needs when they are accused of passing off machine writing as their own. They need to prove a negative. And this system cannot produce one, by construction.
No test can return “a human wrote this.” The only answer available is “not mine.” A clean result from one company rules out that one company. It says nothing about the dozen others, and nothing at all about a model running on someone’s own laptop, which carries no mark and never will. There is no result you can obtain that means innocent.
The accused cannot run the test anyway. A student cannot query Anthropic to clear their name. Neither can a freelancer whose client got suspicious, or a job applicant whose cover letter tripped something. The tool belongs to the company. Access will go to institutions that can sign agreements — the same institutions doing the accusing.
Meanwhile the accusations will keep coming from the bad tools. Schools already own the perplexity-based detectors. They are wired into the grading workflow, they were paid for, and nothing about watermarking removes them. So the machine that generates accusations stays in place, and the machine that could rebut them is locked behind a corporate API. Those are not the same system, and only one of them points at you.
And the middle of the range is where people will get hurt. A middling score does not mean “half written by a machine.” It could mean heavily edited AI text, or a document too short to measure, or mostly human writing with one assisted section, or work that passed through three tools and split its evidence across three keys. All of those produce a similar unremarkable number. Someone will pick a cutoff, and everyone near it will be sorted into guilty or innocent by a threshold they never saw and cannot contest.
The clear cases were never the problem. A fully machine-written document scores high; your own writing sits at chance. It is the enormous ambiguous middle — where nearly all real work lives, because nearly all real work is now partly assisted — where a person gets told they cheated and has no evidence available in the universe to say otherwise.
And writing every word yourself will not protect you. This is the part that should worry people most. The accusations do not come from the watermark — they come from the style-guessing detectors already installed everywhere, and those flag human writing routinely. That Stanford finding is not a hypothetical: more than half of the essays by non-native English writers were called machine-generated, and those were written entirely by hand. You can compose every sentence yourself, on paper, and still be accused. The advice to “just write it yourself” does not actually work, because the thing making the accusation was never measuring whether you wrote it.
Which raises a question worth sitting with: is a number on a scale actually proof of anything?
Not in the way it will be used. A score is a statement about likelihood, not about a person. It says a pattern would be improbable by chance. It does not say who typed the document, whether they wrote it or received it, whether they quoted a source, or whether a collaborator used a tool they knew nothing about. Text has no chain of custody. It can be edited, merged, and pasted, and the number cannot tell any of those stories apart.
There is a name for the mistake that is coming. Statisticians call it the prosecutor’s fallacy: confusing the probability of seeing this evidence if the person is innocent with the probability that the person is innocent. Those are different numbers, often wildly different, and treating one as the other has put innocent people in prison. Every published false-positive rate is an admission that some share of clean documents will trip the test. One percent sounds like rounding error until a university runs a million submissions through it and ten thousand honest people get a letter.
A number is evidence. It was never proof. The distinction will get lost the moment it appears in a box on a screen next to somebody’s name.
The only defense that will actually exist is process. Keep your drafts. Keep the version history, the revision log, the timestamps, the messy earlier draft with the bad paragraph in it. That trail is worth more than any detector result, and it is the one form of proof you control.
So what is it actually for
If it cannot help you check anything, it is fair to ask who it serves. There are honest answers.
Compliance, first and most plainly. The law puts the obligation on the company that built the model. The mark is how Anthropic demonstrates it complied and avoids a fine measured in tens of millions.
Platforms, not people. The intended users are organizations large enough to hold agreements — a social network filtering synthetic political content before an election, a news wire, an app store. Queries at scale, under contract.
Forensics after real harm. When a fabricated document turns up in a court filing or synthetic audio moves a stock price, investigators can ask the vendors. Narrow, but genuine, and worth having.
And one nobody advertises: model makers need to avoid training tomorrow’s models on today’s machine output, which degrades them. A watermark is an excellent filter for that. It may be worth more to the companies commercially than the transparency is worth to the public.
Is it a government control mechanism? Not as built, and I want to be fair about it. The law requires marking, not blocking or reporting. The keys sit with the companies, not with any regulator — nobody handed a government a master key or a registry of who wrote what.
But it is dual-use infrastructure, and that deserves saying out loud. Two years ago, “who generated this text” was a question nobody could answer. Now a handful of companies can answer it on request. Whether that stays a compliance tool depends entirely on who gains the power to compel a query — and subpoenas, licensing conditions, and national-security demands are all ways to compel one. China already requires visible labels, which is a control-shaped posture rather than a compliance-shaped one. The capability is new, it was built fast under a deadline, and almost nobody argued in public about that second use.
Written from an incomplete picture
It is worth being honest about how this rule came to exist, without turning it into a complaint about bureaucrats.
Drafting on the European AI Act began years before most people had used a chatbot. The technology shifted underneath the drafting process repeatedly, and the finished text asks for AI content to be detectable without resolving the question of detectable by whom. That gap is not malice or incompetence. It is what happens when a rule has to be written for a moving target, by people who are necessarily working from a partial view, on a timeline that cannot wait for the picture to fill in.
And the uncomfortable truth is that nobody has the full picture. Not the regulators, not the critics, and not the companies building these systems — who are frequently surprised by their own models and say so publicly. Anyone claiming confident knowledge of where this lands in ten years is guessing. That includes me, and it includes the people who wrote the rule.
I do not think that means regulators should have waited. Waiting for complete understanding of a technology like this means never acting at all, and the harms the rule targets — fabricated evidence, synthetic audio of real people — are already here. Doing something imperfect early is a defensible choice.
But it does mean we should hold this loosely, and expect revision. The first version of a rule written under those conditions is rarely the one that survives. California spent forty years amending its warning law and is still adjusting it. There is no reason to think this will go differently, and good reason to hope somebody revisits the question of who gets to read the mark.
In the meantime, the practical reality is that AI is being folded into the keyboard, the email client, the word processor, the phone, and the search box. Within a few years, producing a document that never touched one of these systems will require deliberate effort. That is not a prediction about the far future; it is a description of the software already shipping. The category the rule is trying to label is quietly becoming the category of everything.
Not the label I expected
I started out expecting Proposition 65.
California passed that law in 1986 with genuinely good intentions: tell people when they are exposed to something harmful. Because the legal risk of omitting a warning dwarfed the cost of adding one, businesses put the notice on nearly everything. It became wallpaper. The state itself now calls the result over-warning and has spent years narrowing the rules trying to make the sign mean something again. Forty years in, they are still at it.
That is not what happened here, and the distinction matters. AI watermarking is not an undifferentiated label slapped on everything. It is carefully proportional. It genuinely tries to measure how much the machine contributed, and it leaves your own writing alone.
It is the inverse of Prop 65. Same disease, opposite symptom. Prop 65 produced a sign everyone can see and nobody reads. This produced a sign nobody is permitted to read. One is meaningless through repetition; the other is meaningless through secrecy. Both are compliance measures wearing the costume of public information.
And I would not call it worthless — that is the easy version of this argument and it does not survive scrutiny. It does real work for regulators, for platforms, for investigators, for the companies’ own training pipelines. It is worthless for the purpose it was announced under. Every genuine use is an argument between institutions: a compliance argument, a legal argument, an evidentiary argument. It was introduced as a tool for public trust. Those are two different products sharing a name.
Where this site stands
I use AI for research, outlining, and editorial suggestions, and I have said so publicly since I started. It does not decide what I recommend, and it does not write the personal parts. I read every article line by line before it publishes and I am accountable for what is in it.
So a watermark tells you something my policy page already tells you, in plain English, voluntarily, in a form you can actually read. That is the difference worth holding onto. A disclosure written by a person can say what the assistance actually was. A mark readable only by its author can only say that it happened, to whoever is allowed to ask.
What I’d do: Keep your drafts — version history, revision logs, timestamps. If you are ever accused, process evidence is the only proof that will exist, and it is the only kind you control. Judge writing by whether it holds up: are the facts sourced, do the links go anywhere real, is a person willing to put their name on it and correct it when it is wrong. And read the disclosure policy of any site you rely on. A written one tells you more than a hidden one ever will.
What I’d skip: Paying for an AI detector. Today’s versions cannot read these watermarks at all, and they carry a documented bias against people who learned English as a second language. Skip assuming that unmarked means human-written — marking is still rolling out, locally-run AI is exempt forever, and heavy editing thins the signal to nothing. And skip accusing anyone based on a detector score. You will be wrong often, and they will have no way to prove it.
Back to the question I asked
When I asked the AI whether it was marking what it wrote for me, it said yes, and then told me it could not turn it off, could not show it to me, and could not confirm it was there. I took that for evasion at first. It was not. It was an accurate description of a system where the thing being disclosed is disclosed to someone else.
The engineering deserves more respect than it is getting. It is proportional, it is careful, and it will not falsely accuse you. But it hands the only copy of the evidence to the companies and the institutions, and leaves the person whose reputation is on the line holding nothing.
So keep your drafts, and keep saying plainly what you did. That has always been the better proof. It is now, quite possibly, the only one you will ever be able to produce.
Verified resources & documentation
- Anthropic — How Claude’s text watermarking works
- Anthropic — How Claude marks AI-generated content
- EU AI Act, Article 50 — transparency obligations
- Regulation (EU) 2024/1689 — full official text
- Google DeepMind — SynthID
- Google — SynthID developer documentation (the machinery, if you want it)
- Liang et al., Stanford — GPT detectors are biased against non-native English writers
- China — Measures for Labeling of AI-Generated Synthetic Content (translation)
- C2PA — the open provenance standard used for images and files
- State of California — About Proposition 65
- OEHHA — short-form warning changes, the state’s own response to over-warning
Keep reading
- How AI is used on this site
- What AI can and can’t do in 2026
- Why you can’t ban AI
- Apple Intelligence and privacy
This article explains publicly documented company and regulatory policy as of August 2026, and simplifies the underlying technology on purpose. Watermarking rollouts are changing quickly, and no public detection tool existed at the time of writing. Nothing here is legal advice.