Could someone train a detector for Claude’s watermark?
This story is from 2026-09-07. It is preserved in the archive; the latest stories are on the live feed.
Anthropic’s Claude watermark appears to work by subtly biasing which tokens Claude chooses using a secret key, rather than adding visible or hidden characters. The idea is to collect a huge number of Claude responses and responses from other models to the same prompts, then train a classifier to te…
Read the full story at r/ClaudeAI ↗
Timeline · 1 report
- 2026-09-07 15:08 · r/ClaudeAI
Could someone train a detector for Claude’s watermark?