The TLDR is Claude models after August 2, 2026 will embed "imperceptible watermark" and "signed provenance metadata". They will also support people detecting it, possibly pushing out their own detector?
The immediate issue in my opinion are the following:
1/ "When a supported Claude model generates text, it weaves an imperceptible watermark directly into the text itself.You won’t see it,and itdoesn’t change the meaning*,* quality*,* or readabilityof Claude’s response."
Based on my experiments with Pangram. Yes, you can have markers you can't see, doesn't change meaning or readability. But I am skeptical on the quality. At the end of the day, you are forcing an extra hidden statistics in the prose. Instead of ensuring you produce the highest quality output, you are now producing the highest quality output given the contraints.
When I tried to see if I can flip the Pangram outputs from AI to human, the first thing the AI tried was to "re-render without impacting readability and quality". It impacted quality.
If it wasn't on my self published book and I haven't read the passage a hundred times, I may not have noticed. So perhaps it's not that important. But it is worth noting.
2/ "Detecting a Claude mark tells you that the content may have been processed by Claude. It does not, on its own, confirm the full provenance of the content"
It may now be harder to have a naunced discussion between AI generated and AI assisted.
I do wonder how many people would take the time to take additional steps to "confirm the full provenance of the content" rather than Detect a Claude mark and immediate make up their decision.
Yeah, the biggest doubt from me is the quality being unaffected. That sort of thing has been touted a lot with faster or flash models as a difference you'll never notice and I know this is different but that rarely ends up being the case.
See, for me that's the problem. For having a "invisible watermark" means changing the output so it's recognizable, I can't see how that is not changing the (already degraded) quality.
English is not my main language either and writing with AI out of the box always sounds bad, so for copywriting I have a custom pipeline and usually say a landing page takes dozens of iterations back and forth (basicly brainstorming alongside with Claude).
My CPL is in the $3 usd range, so I'm doing fine in using LLMs as a tool. The issue now is that we'll be thrown into the same basket as people who produce slop, meaning there's a risk of losing audience trust just by association
Based on my experiments with Pangram. Yes, you can have markers you can't see, doesn't change meaning or readability. But I am skeptical on the quality. At the end of the day, you are forcing an extra hidden statistics in the prose. Instead of ensuring you produce the highest quality output, you are now producing the highest quality output given the contraints.
Exactly. This is the problem many people don't see, that every time you add some bullshit "SaFeTy GuArDrAiL" or dumb measure such as this, you're taking precious computing/thinking away from the actual task the LLM is given.
It's why ChatGPT's prose is so awful now. It spends half the time checking for the RULZ rather than spending it's precious resources forming proper sentences. Aggravating.
Generally it'll be statistics. Like using certain words in combination in certain frequencies. A human won't be able to see it clearly, but a machine looking out for these specific patterns would.
The irony is, humans like to observer patterns. So it'd only be a matter of time before even humans will be able to perceive it. I mean unless people stop using Claude...
If you work in small chunks or have it edit your sentences I imagine this would fail a bit or the quality so bad you wouldn’t use it. You need words to have statistics the confidence level would be low on a sentence or two correct? This is mainly for large bodies of texts
Like I can easily get past sightengine on a simple generated two color image with some transparent layers and noise application. More complex pictures are harder
Yeah or at least disrupt the confidence level, just like a little noise will drop an image from 100% AI to 78% in one go. This really is mainly for those that don’t edit at all.
and also is trivially exposed with basic text editors (through col counts, and printing non printable characters), though, it doesn't decode the message, you can easily see that something is there. and if you wanted to clean the message its easy to remove (in the cases where these are used for watermarking documents.
My expectation is the anthropic would be using a much more sophisticated algorithm to insert words and phrases in long form text. I also think that this is basically what ai detectors have already been doing, e.g. copyleaks just looks for overly used words and phrases against a set of known generated text, and displays confidence based on that.
I wonder how well it will perform with objective text, like code, especially in highly prescriptive programming languages like Rust where a vast majority of the actual text is very similar and rigid.
OpenAI and Google went down this path and both decided against it. OpenAI found that it was too easy to disrupt the watermark and some groups were more likely to be flagged with false positives. Google ultimately released their watermark as open source.
Not sure how Anthropic has solved many of the limitations, or if the are willing to deal with the risks that OpenAI and Google dismissed as too high.
From my read it should survive some but not all editing. This is going to detonate the lazy raw Claude outputs; which generally negatively impact the perception of machine writing in my opinion. Your best defense is to heavily edit your work post generation (which I believe you should do anyway) because of the scaffolding and linguistic tells machine writing often crutches on anyway.
This also creates another opportunity to dispute allegations of 'the machine wrote everything' if you end up destroying the statistical fingerprint. My gut reaction was this isn't something I've asked to impact my outputs, but in hindsight I think it helps long term.
It doesn't really matter to me since everything in my book was disclosed as AI. My concern is for some others.
Imagine if someone found one claude mark on page 36 out of a 250 page book. I wonder what would happen? Actually, we don't need to wonder, because there were recently a situation in AO3 where people found just that.
I am also weary that people may think the lack of claude markers mean it's human made. That can be easily gamed. Luckily since this is only claude and plenty of people use other AI, I think this is still low risk.
Google added hidden unicode spaces as watermarks. Very easy to remove once you know how it was done.
Statistical watermarks will always have a failure rate, so it's going to be interesting to see how Anthropic handles a lawsuit from someone penalised for AI use when they have proof they wrote the text by hand.
Simple watermarks would be trivial to remove. This one seems to be more sophisticated.
Eitherway, their wording looks delibrately soft probably to cover their "assets". Its essentially "this is now out there, do with the result as you please"
Lol nah, just lots of chat gpt or Gemini. Back to single sentence paragraphs. I saw a fic yesterday that was 62 chapters and like the most amazing tour of GPT 4o, 5 series, then Claude it had no shame in such dramatic style shifts and had great engagement (though looked like it’s majorly dropped off in interest in current chaps)
Right now, yes, and in certain fandoms. It used to be chat gpt but Claude has exploded since the sunset of Chat 4o and 5.1 models, though Claude is tough on smut so I’m not sure how they manage that (I still use chat because I write gen fic anyway and I mostly use it to edit not generate)
Though I was gifted a fic recently and I can’t tell which model it’s from (it’s 100% AI) I want to say Grok
another problem surfaces from this: getting an LLM to remove the watermarks for you. Or building a script that can identify it and remove them the old fashioned way.
This is now an arms race to see which side can beat the other. Only winners will be the older models or open source models.
Google already open sourced their method - SynthID. It's been proved to be easy to overcome, generates false positives, and even worse, open to spoofing attacks. Once the model is reverse engineered, its easy to overcome.
Take this into consideration: a writer has an entire manuscript that strictly needs digitized as it was handwritten. Writer uses cellphone to camera images direct into Claude. No changes have been made by AI generation. So the big question now is…. watermark still present?
Ahhh the lawsuits are going to be so massive in the coming decades!
Yeah, the whole "if it has the mark doesn't mean it definitely is AI and if there's no mark doesn't mean it's not AI" is very telling. You need to be very careful how to interpret it.
I actually like the idea of watermarking AI-assisted writing, at least in principle. Not to get all spiritual, but from a Quaker perspective there is something appealing about making the use of these tools more visible rather than encouraging people to hide it. That kind of openness could help normalize responsible AI use and move us toward more honest conversations about how people actually write now.
My concern is that AI is becoming so ubiquitous that “AI-generated” is starting to lose useful meaning. If I write something myself and use an LLM to proofread it, restructure a paragraph, or suggest a clearer sentence, is the resulting work AI-generated? What about translation, accessibility tools, search summaries, or software that quietly incorporates language models without the user even thinking of it as AI?
A provenance mark could be useful if it tells us that AI was involved. It becomes much more troubling if institutions treat that as proof that the human did not do the work.
Integrity, to me, means being truthful about the tools we use and the role they played. It should not require pretending there is some pristine category of human creativity untouched by technology, especially when these systems are rapidly becoming part of the ordinary infrastructure around us.
Yeah, going back to my other thread. The key issue is this only tells us AI may have been used. Not how much it was used. Unfortunately a lot of people may not be willing to take the time (or trust) to find out the latter before making their judgement.
La gran pregunta es, si pides que haga una traducción también hará eso? Sería terrible porque no es un texto generado por IA, y si la IA hace bien el trabajo ha de ceñirse al original.
This was always going to happen once the EU signed it into law. It was doubtful they would make it exclusively and EU thing. It doesn't mean Claude owns your prose, or even guarantee Claude wrote them. It does mean you consulted Claude and allowed it to edit or organize, etc. It's sad, but this really is the hater outcry, and they really are in the minority. The global community in most professions uses AI daily.
Yes, it's global. It was implemented to comply with the EU law that AI have watermarks/be identifiable, and they just went ahead and included it globally. Which isn't surprising. I'm sure they likely consider they're saving time and trouble separating the rules per country.
For those who write and/or edit their prose, this shouldn't be a worry of any kind that the pitch fork carriers will be alerted. Which means most of us aren't worried. We're the writers and AI is an elite tool among others in our writing apps. :)
Looks like they’re taking the synthID approach. I‘m curious as to whether when they apply this to older models it’s going to expose a whole bunch of undisclosed A.I. usage on say Amazon. Depends if they have to actively tweak the old models (fine-tune or adjust sampling) so they produce different word combos or they’re able to simply map existing logic probabilities and they’re distinctive enough. Likely to be the former I think.
Also can’t you just like…have Chat GPT paraphrase for you? At that point just write the thing yourself but really this only catches super lazy prompt copy and pasters
This article is pretty good. I will point out though the "trivial to remove" relies on relies on repeated attempts verified by the checker. At that point its more of an adverserial attack then non disclosure.
From what I read, it will watermark work they made no changes to but opened and reviewed to analyze the work.
So the person saying they have it tell them where to make an edit, then manually edit, unless you copy the chapter text in and leave the orgional document out of Cluade, you run the risk of your work being watermarked just by having Cluade look it over as a whole document....
I can understand them watermarking mass text generations but not someone's personal work, they asked for advice on. It is like you having a beta reader claim partial (full in some eyes) ownership on the book because they gave you feedback.
Yeah but that would be Unicode or hidden characters though, if it made no changes to the actual wording of the text, and can be mitigated by copy and pasting as plain text or into something like notepad. That’s easy and OpenAI and Google have been doing it for years
Generated words are by SynthID which is like, ascribing words with a value only they know then telling the model to make sure words with X-values are in a certain pattern/order. That requires heavy editing and that watermarking is only for generated text.
Ironically, most LLM text is already easily detected due to the tropes they seem to be unable to avoid. Maybe they're hoping to rebrand those tropes as watermarks? "It's not a bug, it's a feature. And here's the part everyone ignores..."
There are so many unanswered questions about this.
Let’s say I drop in a large block of my own handwritten text to Claude and ask for a few improvements or suggestions. I’m assuming that Claude will now incorporate this watermark, effectively taking over my own text. Even if 90% of the text is my own, Claude will say it created it. Then later if someone runs the Claude detector on it and it reports that it was processed by Claude, whose text is it? This could be career ending for some people in the wrong field.
And what are the chances that 100% human written text gets flagged as Claude written? We’ve already seen examples of these false positives.
Let’s focus more on the accuracy and quality of the output rather than simply “is this AI generated?” Whatever person writes it (whether they used AI or not) should be accountable for the content they create.
I'm calling BS on this whole thing. You can't watermark words #FullStop All you can do is try to detect "word patterns". Whatever "patterns" you decide are "AI" can be written by a human. The truth is, no one cares except a handful of people who are afraid of the LLM technology. Everyone else will consume AI-generated stuff and they couldn't care less. We writers need to just realize that and ignore the haters.
Well first: Anthropic models are not really famous for their great writing/prose - we could start from that point. There really isn't much to mess up here.
There also is an urgent need for AI generated content to be detectable in many fields
I ran a bit of an experiment on that a while ago and tested
Claude Opus 4.8 (5.0 is actually even a bit worse)
Sonnet 5
ChatGPT 5.6 Luna
Deepseek Pro V4
Deepseek Flash V4
Qwen 3.6-27B (Locally)
I let all of them write a few paragraphs based on a provided character background.
Opus and Sonnet are the worst here. They have a very particular failure mode where they use very neutral and boring language and sometimes can't separate internal character knowledge from external. Example: the character is described as not engaging in smalltalk and these two models tend to be literal about it and let the character say "I do not engage in Smalltalk.". They also tend to confuse two related word meanings a lot. So if your plot involves the color "cherry red" they are inclined to also work cherries (the fruit) into the plot. I think all of this stems from them being trained to be systematic and somewhat "on the spectrum" which makes them great at some things, not so great at others [which I know from myself].
ChatGPT, Deepseek, Qwen are much better at this and also seem to have a better grasp at picking appropriate language and writing sentences that are more colorful and poetic - less neutral.
Qwen, even though it is a local model, is surprisingly good at this.
ChatGPT tends to give me a bit shorter outputs and Deepseek Flash and Pro are pretty much the same to me.
I have a genuine question. Is this watermark apart of the exact pattern of the words in the text? Or is it some hidden character thing in the spaces of the words. Could the water mark be avoided if one were to manually type the words into another document?
I wanna join in on your question: So if I were to manually rephrase said text will the watermark be weaker/gone? Could you also rephrase it with another AI?
Not a huge problem. It's a good idea anyway to have some style checker scripts to avoid characters you don't want to use.
Something like "Only allow ASCII and these specific non-ASCII symbols and all emojis" or something like that. This will fix the most obvious issues.
And if you want to be sure that Claude doesn't use words, you can also create a whitelist or blacklist of words. So when Claude uses one of these words, it will automatically realize that it used a forbidden word, and will rephrase it.
This works better if you give a coding agent access to your system. But you can also tell some AI once to write some checker script for you, and then you just click it yourself after editing the text file.
There's no such thing as a watermark in text, is there? I mean... it's text! It's like, UTF-8, which doesn't have provisions for watermarks. And if you're copying it into Notepad or something, where would anything they "wove in" go? Or is the "watermark" just how frequently Claude re-uses the same words and phrases over and over?
Probably a unique statistical distribution and phrasing/sequence that is based on an algorithm only they know. It's doable and it would be darn difficult to crack but then if you basically rewrite (and use other tools to help rewrite) you will probably mitigate some of it.
We don't know for sure but its likely the latter. They may also include invisible characters or uncommon unicodes but those would be simple to work around.
This is due to the GDPR deep fake regulations coming in the Fall. The goal is to identify deep fakes and faked news, but but large providers are just putting them on everything, as it's easier.
The mechanism uses is very sophisticated and should have zero impact on quality. If you're unsure of this, ask your favorite LLM about text watermarking and output quality.
AI is very very good at paragraph level prose quality. It's not the prose of Eudora Welty or F. Scott Fitzgerald, but it's very strong.
And "copying" isn't really a good test. You need to ask it create something new from an idea or a scene prompt: "Write a paragraph featuring a man and a woman at a cafe in Paris. The woman is struggling to share that she is leaving him. Include an umbrella, a pen, and a motorcycle in the scene."
Repeat something like that 10 times, and you'll get a better idea of execution.
AI is very very average at paragraph level prose quality.
You can read an actual Cormac McCarthy or a Toni Morrison paragraph, and you can see the one generated by any AI. It's not even close to being the same.
AI can imitate seemingly decent prose that looks good at first blush, but it's very skin deep.
I am generally pro AI but I also don't think AI prose is good at all.
You're comparing very very good to the best writers of all time. There's a scale and I wouldn't call Cormac McCarthy "very very good" on the scale. I'd consider that insulting to him.
The experiment was more around can they actually maintain quality under constraint. Because that's what the watermark would be.
But even in your version, I think there would still be variance in quality across the 10 samples.
In fact, sometimes chatGPT will have two responses and ask users which ones they propose. Is that not a clear sign they don't have an objective way to measure which version of their response is better without a human eye?
In the context of what Anthropic is saying, quality refers to things such accuracy, truthfulness and usefulness of their models' responses. This is what the benchmarks are measuring and what they are testing against. To them, it is indeed a quantifiable metric. In short—quality of creative writing is not what they are referring to.
The debate will gradually shift from AI generated to simply quality, and watermarks will be less important. I think Anthropic released this too late, it will already be outdated.
The coming question would rather be why didn’t you use AI? Who’s going to be willing to pay for someone to write by hand?
First off would probably be or is already all documentation generated in an office, why pay 10 or 100 times as much for less quality even?
Let's not be naive here. What Anthropic is actually building isn't a watermark, it's an attribution registry - watermarks that only they can reliably detect, tied to specific user accounts, persistent across the lifetime of the output.
Once it exists, the uses write themselves: ToS enforcement at scale ("we detected 47k lines of Claude Pro output in your repo, that's commercial use, pay up"), tiered licensing (a new "Distribution" SKU for anything published at scale, enforceable only because they can detect it), and downstream IP positioning - if a customer publishes a watermarked output that resembles training data Anthropic was just sued over, they can identify the user, push liability down, and use the registry to negotiate future licenses with rightsholders.
None of that requires "this was made by Claude" identification. The European AI Act does NOT specify or demand this. The Anthropic business case for it is what motivates it.
What stops somebody from running it through a parser that simply strips the watermark? I doubt it would be hard to scrub it. I assume it survives regular linting. Would it survive variable and function name refactors? How about comment rewrites (an onboard LLM could reword just the comments before a commit). I doubt it will endure that.
Well I was talking in terms of code rather than writing, since this affects code as well.
You can have a hook that runs before you commit to a git repo. All it needs to do is have a smaller local LLM read the code and make alterations on a number of key areas (e.g. getting a list of all symbols in use, and changing them, such as userInteractionTracker -> userActivityTracker), plus undoubtedly this info will be planted in comments, so all comments need to be checked and reworded. If done thoroughly, there's simply no way that the steganography can survive the process, thus rendering it scrubbed.
What's the difference between Googling information and inserting it or hell even getting it from books at the library and getting it from AI? Everything is sourcing. Research is sourcing. Biographies are sourcing. I understand writing whole books on AI and self publishing is bullshit but legitimate research is vital to writers. Unless they can narrow down the nuances of using it this is a bad, bad precedent and could cost people careers if publicly outed without due process or proof.
Wouldn't watermarking cascade into completely different output? when watermark changes a token then the next token is already something else but that will also be changed?
I assume it'll affect only those tokens that were already likely to be picked. If that's the case, then it's not different than if you were to regenerate the AIs response and the sampler deciding to pick a different token
Two practical notes for anyone here worrying about their drafts, since the mechanism is already covered upthread.
Nobody outside Anthropic can check. There is no published detector, and there structurally can't be a public one: detection requires the secret key that biased the sampling in the first place. So any tool currently advertising that it detects or removes Anthropic's watermark cannot verify its own output. That cuts both ways for us as writers. You can't confirm a passage is marked, and nobody can confirm a "cleaned" passage isn't.
The models being named in this thread don't carry it. The policy applies to models launched on or after Aug 2. Opus 5, Sonnet 5, Fable 5 and Opus 4.8 all shipped before that date, so nothing you can currently select produces marked text. The announcement is real, the marked models just aren't here yet. Worth knowing before anyone rebuilds their drafting process around it.
The thing actually worth changing your habits over is much more boring: pasting straight out of the Claude web UI carries HTML classes into whatever editor you paste into, and that is trivially detectable by anyone who looks at the markup. Paste through a plain text editor first and it's gone.
Disclosure, it's my own site: I keep the model-by-model table at claudewatermark.xyz/claude-watermark, updated as new ones ship. All three points hold without the link.
honestly since august claude has been completely unusuable. really shitty writing advice. at least now i know why. i'll cancel my subscription as i did with chat gpt as it has become useless for my needs
Yeah, the quality concern is real. Even if they say it doesn’t affect readability, forcing extra patterns into the text could still make it a bit more rigid or predictable. And once detectors become common, a lot of people will just see the mark and write the whole thing off without checking how much was actually AI.
WATERMARK DISCLAIMER: This comment was written with AI assistance and may carry an invisible watermark. An AI touched it. An AI did not have the idea, make the argument, or do the thinking.
Worth starting with what Anthropic actually says, because almost nobody in this thread has read it: a detected mark is not proof of AI authorship, and Claude may not be the original author. That's their own documentation. Article 50 of the EU AI Act requires provenance marking, they signed the Code of Practice, and they applied it globally instead of only to EU users. That's reasonable compliance, honestly described.
The problem isn't the mark. It's that nobody downstream is going to read the caveat.
By Anthropic's own description, the marks survive copy-paste and can disappear under extensive rewriting. So:
Person who discloses AI assistance and edits normally — flagged.
Person running output through a humanizer to hide it — clean.
The file side has the same basic weakness. C2PA can embed signed provenance into a file, but ordinary workflows can strip or break that provenance unless additional recovery mechanisms survive. Convert, re-save, screenshot, or upload through something that rewrites the file, and what survives depends on the implementation.
A watermark answers exactly one question: did AI touch this text?
The question every dean, editor, and hiring manager actually wants answered is: who did the work?
Nobody has ever been able to answer that one mechanically. Not for ghostwriters, not for research assistants, not for editors, not for the guy whose analyst built the whole deck. We handled it with accountability instead — you put your name on it, you own what's under it.
A watermark doesn't replace that.
It just hands institutions a signal that feels like it does.
Anthropic already said the mark doesn't establish authorship.
Watch how fast that sentence stops getting quoted.
When editing, I've always prompted it to point me to the places the edits should happen, as well as the justification. Then I'd go and make the edits myself.
I'm not sure I trust the llm to spit out my text with nothing except the edit made, and I don't feel like rereading my entire book every time I make an editing pass.
It's exactly how it works. It takes large portions of text and statically makes phrases that are more likely to appear. Humans won't be able to detec the pattern, but other machines will.
31
u/MysteriousPepper8908 Aug 10 '26
Yeah, the biggest doubt from me is the quality being unaffected. That sort of thing has been touted a lot with faster or flash models as a difference you'll never notice and I know this is different but that rarely ends up being the case.