Anthropic just added an invisible watermark to everything Claude writes — but how does hiding a signal inside text even work?
When Claude generates text, the watermark isn't a visible tag — it's baked into the token selection process itself. The model subtly biases word choices in ways imperceptible to readers but detectable by statistical analysis. Anthropic hasn't published its exact method, but this distributional watermark approach is the leading technique in the field.
The catch is fragility. The signal can vanish if text is heavily edited or paraphrased, and it can't prove authorship: paste your own essay into Claude for a quick proofread and the output may carry a Claude mark even though the ideas are entirely yours. Unmarked text isn't proof a human wrote it, either.
For images and files, Anthropic uses digitally signed C2PA provenance metadata embedded in the file — though it disappears if the file is screenshotted or converted.
Two tools, two different failure modes — neither is a reliable AI detector on its own.