Anthropic’s new Claude watermarking system is exposing a growing problem with AI authorship: detecting that artificial intelligence participated in creating a piece of text does not necessarily tell us whether the ideas, writing or final work actually came from a machine.
WHAT’S HAPPENING
Anthropic is implementing invisible text watermarking in Claude as major AI providers respond to transparency requirements under the European Union’s AI Act.
Unlike a visible label saying “AI generated,” the watermark is embedded statistically in the words Claude chooses while producing text.
Anthropic says its approach is based on Google DeepMind’s SynthID-Text technology. Rather than adding hidden characters or metadata to ordinary text, the system subtly changes the source of randomness Claude uses when choosing between equally reasonable words.
Over a sufficiently long piece of writing, those choices can form a detectable pattern. (Anthropic)
But there is an important limitation.
Anthropic says detecting the watermark can indicate that Claude was likely involved with the content.
It cannot determine whether:
Claude originally wrote it
or
a human wrote it and Claude substantially edited it.
Translations created by Claude can also carry the watermark because Claude selects the words of the translated version. Heavy editing can create the same issue. (Anthropic)
Light proofreading is different. If Claude merely fixes a handful of grammatical or punctuation errors, Anthropic says there may be too little AI-generated language for the watermark to register reliably. (Anthropic)
That means the watermark is really evidence of AI involvement, not necessarily AI authorship.
WHY IT MATTERS
For years, the debate around AI-generated content has largely been treated as binary:
Human written or AI written.
That distinction is becoming harder to maintain.
A person might research and write an entire article, then ask AI to reorganize paragraphs, tighten sentences or translate it.
Another person might ask AI to produce almost the entire first draft and then extensively review, rewrite and approve the final version.
Those are very different creative processes.
Yet a simple AI label may struggle to explain the difference.
The European Union’s own transparency rules illustrate the problem.
Article 50 generally requires organizations publishing AI-generated or manipulated text about matters of public interest to disclose that AI was involved.
But the law provides an exception when the material has undergone human review or editorial control and a person or organization accepts editorial responsibility for its publication. (EUR-Lex)
At the same time, AI providers face a separate requirement to make generated material machine-detectable where technically feasible.
Those two requirements are measuring different things.
One asks:
Was AI involved?
The other increasingly asks:
Who reviewed the information and accepts responsibility for publishing it?
That distinction could become enormously important.
WHO BENEFITS
Publishers and platforms gain another tool for identifying when generative AI may have participated in producing content.
Schools and universities could potentially gain better evidence when investigating undisclosed AI assistance — although Anthropic itself warns that a watermark cannot establish who actually authored the underlying work.
AI developers gain a technical mechanism for complying with emerging transparency regulations without visibly altering every piece of generated text.
Consumers could eventually gain greater transparency about where AI has participated in producing the information they encounter.
But the usefulness depends heavily on people understanding what the signal actually means.
WHO LOSES
Anyone who treats an AI watermark as proof of cheating or machine authorship could reach the wrong conclusion.
A writer who produces original work but substantially edits or translates it through Claude could generate a detectable watermark.
At the same time, the opposite problem exists.
The absence of a watermark does not prove something was written by a human.
Anthropic says detection becomes weaker with short passages, factual material where fewer word choices are available, older models and text that has been substantially rewritten. (Anthropic)
Developers have also already begun experimenting with tools designed to weaken or remove Claude’s watermark through rewriting, highlighting how difficult universal AI detection may become. (TNW)
So both conclusions can be dangerous:
Watermark detected ≠ AI authored everything.
No watermark detected ≠ human authored everything.
WHAT HAPPENS NEXT
The biggest challenge may not be building better AI detectors.
It may be deciding what we actually want to detect.
As AI becomes embedded in writing, research, translation, editing, coding and everyday professional work, the line between human-created and AI-created content will continue to blur.
Schools may need to distinguish between prohibited AI authorship and permitted AI assistance.
Publishers may need to distinguish between machine-generated material and human-controlled work created with AI tools.
Employers may need policies determining how much AI participation matters when evaluating someone’s work.
And regulators may increasingly focus less on whether AI touched a document and more on who ultimately takes responsibility for its accuracy and meaning.
Anthropic’s watermark is therefore exposing a much larger issue than whether Claude-generated text can be detected.
The technology can potentially tell us that AI participated.
It still cannot answer the more important question: who actually created the work?