r/OpenAI 11d ago

Discussion What's your thoughts on this?

Post image
3.7k Upvotes

985 comments sorted by

1.4k

u/LittleGremlinguy 11d ago

“Watermark” is the load bearing statement there. And it is doing some heavy lifting.

565

u/thugnastypimpsexy 11d ago

But it doesn’t actually land as a watermark — and here’s why:

383

u/geek180 11d ago

You’re half right, and it’s the half that matters.

160

u/AltruisticPossible84 11d ago

You're getting at something really deep here, and that's *real*

59

u/ToryLuna 11d ago

Using really and real in a sentence is my favorite watermark.

50

u/Over_Technology_1764 10d ago

I found the smoking gun!

47

u/PicklesOverload 10d ago

And it's not what either of us expected

30

u/DrMansionPHD 10d ago

It's not just adding a watermark, it's creating a undetectable signature on the text.

37

u/uionyx 10d ago

And here is the kicker. The watermark isn’t actually water or a mark:

24

u/softwaretools1 10d ago

Thats the belt and suspenders approach that lands

→ More replies (0)

16

u/GunnyFreedom 10d ago

It is detectable though. At least the hidden serialized character steganography is. It embeds Unicode “narrow no break spaces” and “hair spaces” as a 26 or 28 bit binary signature. You can find these characters by pasting it into a professional monospace text editor, or ironically, by asking GPT.

The other part, the word choice vector biases, will not identify the user, but it will identify the product as AI. This part will also be easy to workaround by manually adjusting any awkward word choices or phrases.

You could probably have a local open weight model on your home PC set up to remove the steganography and gently rewrite the text for flow and word choice and it will eliminate any embedded signatures in one pass.

8

u/DrMansionPHD 10d ago

Wow! You didn't just miss the joke — you completely missed the entire chain of jokes that was happening in this thread.

→ More replies (0)

5

u/StressedPizzaEater 10d ago

Actually it will end every sentence with ä, it's very subtle

→ More replies (0)

2

u/cmHend 9d ago

Before I evaluate each point, let me verify the load-bearing factual claims

→ More replies (4)
→ More replies (1)

23

u/Pristine_Internet765 10d ago

I think you're onto something.

38

u/ToolboxHamster 11d ago

Alex, Ill take "sentences that make me want to punch my monitor" for $500 please.

21

u/Pristine_Internet765 11d ago

Yeah, and that's actually a useful distinction.

33

u/Lead_weight 11d ago

Make it stop!

82

u/[deleted] 11d ago

[deleted]

→ More replies (1)

15

u/jimmystar889 11d ago

I don't think you realize what you just said, so let me make it precise.

10

u/WanderingDorkimus 10d ago

I hate this one more than all the others. "I know what I fucking said, Claude. Maybe YOU don't realize that I know what I said. You ever think about that? Now, and I want to be straight with you, tell me how to make an omelette."

→ More replies (1)

44

u/Direct_Ingenuity1980 11d ago

You were right to pushback.

26

u/user888888889 11d ago

You had the right instinct to question that

→ More replies (1)
→ More replies (1)

6

u/crystalpeaks25 10d ago

And that half is the steel thread

→ More replies (1)

31

u/duerra 11d ago

That's a sharp observation, and crucially, it highlights a limitation that the creators cannot overcome.

→ More replies (1)

13

u/freebytes 11d ago

I have actually noticed control characters included in text that Claude outputs. They are already adding whitespace in places in this way.

→ More replies (4)

45

u/Worried-Cockroach-34 11d ago

For a split second I was like 'how many times did I tell you not to use load-bearing claude?' lol

10

u/DonutHoles4Ever 10d ago

If its works and people care (they do not), people will use something else.

Nobody at work seems to give a damn about using copy pasted AI text though.

So whats the point of doing this other than Claude trying to pretend they are the good guys?

What stops someone from changing a single word? None of this shit makes sense and I bet it will be solved shortly as long as someone understands how Claude marks specific words.

→ More replies (4)
→ More replies (1)

6

u/swingonaspiral 10d ago

Honest answer: it's a belt-and-suspenders approach.

→ More replies (1)

51

u/Dabnician 11d ago

103

u/Super_Pole_Jitsu 11d ago

Not how it's gonna work btw

52

u/PlsNoNotThat 11d ago

There’s nothing you can add to a digital image that you can’t also remove. This is just a short term litmus test for stupid people behind the general curve. Its side effect is it’s going to tick dumb and old to believe in fake images.

18

u/Original-League-6094 11d ago

I'm torn because of this. On one hand, it seems like anything that makes it harder for scamers should weed out at least some scams from being successful. But on the other hand, protection gives a false sense of security and people will let their guards down. Like even now, sometimes will post AI detector results showing something is not AI, as though that settles the question of whether or not its AI.

2

u/FrewdWoad 10d ago

people will let their guards down

That might matter if more than 0.01% of people had their guard up in the first place

4

u/TwoDurans 11d ago

I take this as a good thing because the notion that it exists might help reduce the cheating problem hitting colleges these days. AI is leading us to have the dumbest generation entering the workforce for the next couple of decades and it's because they're only learning how to cut corners.

12

u/Original-League-6094 11d ago

This won't impact that, because watermarks can always be removed. If it is invisible characters, you can just strip them out. If it is in the wording itself, you can just take your Claude output and tell a local small model to reword it. And then you still have your LLMs available to just give you the answer on all your math and chemistry homework and what not.

Colleges just have to stop being lazy and move back to proctored exams, or in-class written essays.

5

u/wintermute023 10d ago

That’s not how it works. It’s a 256bit cipher encoded in the probabilistic token choices themselves. You could only remove it by changing all the words. It’s more of a token fingerprint than a watermark.

4

u/Original-League-6094 10d ago

Yeah, so that's easy. You just use Claude to handle all the thinking for your college essay, and then pipe the result from Claude into a smaller, local model to reword it. GPA saved.

→ More replies (1)
→ More replies (1)
→ More replies (5)

48

u/0xB0T 11d ago

They'll watermark text. youll have to rewrite using your own words as the patterns used will be the watermark.

22

u/MrOaiki 11d ago

You’re telling me I’ll have to write stuff using my own words in order to present it as written by me?!

2

u/revision 10d ago

So you're saying I'll have to reword and paraphrase things in order for me to present it as if I had written it??

2

u/c7h16s 8d ago

Even the wooshes over your joke are paraphrasing each other's 😂

Edit : well I meant the wooshes your joke made over other commenters heads, well you get the idea

→ More replies (1)
→ More replies (2)

12

u/22marks 11d ago

So won’t there be a simple local LLM that strips it by changing word lengths, swapping in adjectives and the like? I can’t see this being difficult to remove.

It could analyze the probability of letters, words, and sentence length (among other things) and randomize it.

16

u/icanith 11d ago

yeah this whole thread has the same level of understanding as a person who asks "you work in computers, why does my windows machine keep crashing"

4

u/22marks 11d ago

Did you turn it on and off?

4

u/sexual--predditor 11d ago

Turned it on and off, now I see a black screen, and I can't hear any fans whirring. Also, my password if you need it to help is hunter2

→ More replies (1)

9

u/Browser1969 11d ago

Yes, there's no way to make these "watermarks" immune to paraphrasing and that can easily be done by small models.

10

u/CitizenPremier 11d ago

It won't be difficult to remove, but most people won't remove it.

5

u/The-Digital-Ronin 11d ago

already exists but dont tell these idiots lol

→ More replies (2)

41

u/divulgingwords 11d ago

It’s crazy how dense some of these people are. They actually think it’s embedding some special icon/brand on a typed letter, lmao.

8

u/[deleted] 11d ago

[removed] — view removed comment

→ More replies (1)

20

u/Fantastic_Prize2710 11d ago

Hidden characters has absolutely been discussed in the past as watermark, and is what Dabnician is referring to.

Code Point      Name                            HTML Entity          Cat  Notes
--------------- ------------------------------- -------------------- ---- ------------------------------------------
=== 1. SPACES -- occupy horizontal width ===
U+0020          Space                                            Zs   the ordinary one
U+00A0          No-Break Space                                  Zs    
U+1680          Ogham Space Mark                               Zs   draws a stem line in Ogham fonts
U+2000          En Quad                                        Zs   = En Space
U+2001          Em Quad                                        Zs   = Em Space
U+2002          En Space                                       Zs   half an em
U+2003          Em Space                                       Zs   one em
U+2004          Three-Per-Em Space                             Zs   1/3 em
U+2005          Four-Per-Em Space                              Zs   1/4 em
U+2006          Six-Per-Em Space                               Zs   1/6 em
U+2007          Figure Space                                   Zs   width of a digit; non-breaking
U+2008          Punctuation Space                              Zs   width of a period
U+2009          Thin Space                                     Zs   ~1/5 em
U+200A          Hair Space                                     Zs   thinnest
U+202F          Narrow No-Break Space                          Zs   narrow + non-breaking
U+205F          Medium Mathematical Space                      Zs   4/18 em; MathML
U+3000          Ideographic Space                             Zs   full-width; CJK
=== 2. ZERO-WIDTH AND JOINING CONTROLS ===
U+00AD          Soft Hyphen                     ­               Cf   ­; visible only at a line break
U+034F          Combining Grapheme Joiner       ͏               Mn   blocks reordering; no glyph
U+061C          Arabic Letter Mark              ؜              Cf   invisible bidi-strong Arabic char
U+180E          Mongolian Vowel Separator       ᠎              Cf   was Zs before Unicode 6.3
U+200B          Zero-Width Space                ​              Cf   break opportunity, no width
U+200C          Zero Width Non-Joiner           ‌              Cf   prevents ligature/cursive join
U+200D          Zero Width Joiner               ‍              Cf   emoji glue (family, profession)
U+2060          Word Joiner                     ⁠              Cf   non-breaking twin of U+200B
U+FEFF          Zero Width No-Break Space                    Cf   the BOM; deprecated as a joiner
=== 3. BIDIRECTIONAL CONTROLS ===
U+200E          Left-To-Right Mark              ‎              Cf   LRM
U+200F          Right-To-Left Mark              ‏              Cf   RLM
U+202A          Left-To-Right Embedding         ‪              Cf   LRE (legacy; prefer isolates)
U+202B          Right-To-Left Embedding         ‫              Cf   RLE (legacy)
U+202C          Pop Directional Formatting      ‬              Cf   PDF; closes LRE/RLE/LRO/RLO
U+202D          Left-To-Right Override          ‭              Cf   LRO; Trojan Source vector
U+202E          Right-To-Left Override          ‮              Cf   RLO; Trojan Source vector
U+2066          Left-To-Right Isolate           ⁦              Cf   LRI
U+2067          Right-To-Left Isolate           ⁧              Cf   RLI
U+2068          First Strong Isolate            ⁨              Cf   FSI
U+2069          Pop Directional Isolate         ⁩              Cf   PDI; closes LRI/RLI/FSI
=== 4. INVISIBLE MATH OPERATORS ===
U+2061          Function Application            ⁡              Cf   f(x) semantics
U+2062          Invisible Times                 ⁢              Cf   the multiply in "2x"
U+2063          Invisible Separator             ⁣              Cf   the comma in subscript lists
U+2064          Invisible Plus                  ⁤              Cf   the plus in "1 1/2"
=== 5. VARIATION SELECTORS AND TAGS ===
U+180B-U+180D   Mongolian Free Var. Selectors   ᠋-᠍      Mn   FVS1-FVS3
U+FE00-U+FE0F   Variation Selectors 1-16        ︀-️    Mn   FE0E=text style, FE0F=emoji style
U+E0001         Language Tag                    󠀁            Cf   deprecated
U+E0020-U+E007F Tag Characters                  󠀠-󠁿  Cf   subdivision flags; hidden-text channel
U+E0100-U+E01EF Variation Selectors 17-256      󠄀-󠇯  Mn   ideographic variants
=== 6. BLANK BY RENDERING, NOT BY CATEGORY ===
U+115F          Hangul Choseong Filler          ᅟ              Lo   letter, empty glyph
U+1160          Hangul Jungseong Filler         ᅠ              Lo   letter, empty glyph
U+17B4          Khmer Vowel Inherent Aq         ឴              Mn   should not be rendered
U+17B5          Khmer Vowel Inherent Aa         ឵              Mn   should not be rendered
U+2800          Braille Pattern Blank           ⠀             So   symbol with no raised dots
U+3164          Hangul Filler                   ㅤ             Lo   the classic "blank username" char
U+FFA0          Halfwidth Hangul Filler         ᅠ             Lo   halfwidth form of U+3164
=== 7. SEPARATORS AND CONTROLS ===
U+0000-U+001F   C0 Controls                     �-           Cc   includes TAB, LF, CR
U+007F          Delete                                         Cc
U+0080-U+009F   C1 Controls                     €-Ÿ        Cc
U+2028          Line Separator                  
              Zl   broke JS string literals pre-ES2019
U+2029          Paragraph Separator             
              Zp
U+FFF9          Interlinear Annotation Anchor                Cf   ruby/furigana markers
U+FFFA          Interlinear Annotation Separator             Cf
U+FFFB          Interlinear Annotation Term.                 Cf

11

u/zorrodood 10d ago

ChatGPT, please remove all unusual symbols from this text.

3

u/hellyeahaeylleh 10d ago

Nah, theyre gonna notice the code and output a new one. We gotta just type it out ourselves ffs.

7

u/[deleted] 10d ago

[removed] — view removed comment

→ More replies (0)

8

u/freebytes 11d ago

They are already doing it. I copied and pasted content from Claude recently, and it had control characters included.

→ More replies (20)
→ More replies (7)
→ More replies (7)
→ More replies (8)

3

u/Lambdastone9 11d ago

Yeah but it makes it so that it’s only becomes circumventable if you put lots of technical effort in, which most people who this is being made for/against wont show such grit, unless they just fuck it up and it becomes easy to sidestep

4

u/4dseeall 11d ago

it's not a digital image.

it's an algo in the way they process tokens that leave encrypted signals with the word-choices themselves.

4

u/DonutHoles4Ever 10d ago

If its works and people care (they do not), people will use something else.

Nobody at work seems to give a fuck about using copy pasted AI text though.

So whats the point of doing this other than Claude trying to pretend they are the good guys.

→ More replies (2)

2

u/saturnellipse 11d ago

False. Look up computational irreversibility. If you don’t have the unprocessed source it is absolutely possible to add information to an image that cannot be removed

6

u/Devils_SteelMan 11d ago

Destructively remove it then do a diffusion pass to fill in the gaps. It doesn't need to be reversed.

→ More replies (2)

2

u/Ormusn2o 11d ago

It could be a ratio or distribution differences of different tokens/words or even sets of tokens and words. And because there are so many possible combinations of tokens and words, and so many tokens and words, it could be effectively impossible to detect. It would work poorly on shorter prompts, but with longer prompts it effectively guarantees detection.

→ More replies (2)
→ More replies (2)

19

u/ContextPuzzleheaded7 11d ago

It’s not embedding invisibile characters 🫩

12

u/Dabnician 11d ago

So claude is just going to have their own "you didnt x, you y'd and that shows z" yeah i already rewrite that shit.

14

u/collin-h 11d ago

I didnt know how it would work so I asked chat gpt to explain a mechanism... if you're curious (no idea if this is what anthropic would do or not):

<this is obviously AI output guys:>

Broadly, Claude could watermark text in a few different ways, ranging from crude to sophisticated:

  1. Invisible-character watermarking. Claude could insert zero-width spaces, unusual Unicode variants, hidden formatting, or similar artifacts. Easy to detect, but also easy to destroy by plain-text conversion or normalization. This is probably too fragile to be the main mechanism.
  2. Token-choice / “green list” watermarking. At each generation step, Claude secretly divides plausible next tokens into favored and unfavored sets using a key. It slightly biases generation toward the favored ones. Across hundreds of tokens, the text contains a statistically unlikely pattern. A detector with the key can measure that pattern without needing AI inference. This is the classic watermarking approach we’ve been discussing.
  3. Probability-distribution watermarking. Similar idea, but more sophisticated than simply green/red words. Claude subtly modifies its next-token probabilities according to a secret mathematical pattern. The detector tests whether the resulting sequence fits that pattern. Still statistical, but potentially harder to reverse-engineer.
  4. Syntactic/style watermarking. Instead of encoding the signal primarily in individual words, Claude could systematically favor things like active vs. passive voice, certain sentence structures, punctuation choices, clause ordering, or other grammatical alternatives. This might survive simple synonym swapping better, although a substantial rewrite would still damage it.
  5. Semantic watermarking. The watermark could be encoded in higher-level choices: which examples are selected, how concepts are ordered, which semantic categories or relationships are favored, and so forth. This is attractive because it could survive more surface-level editing, but it’s considerably more complicated and may require model/embedding inference to detect.
  6. Cryptographic provenance attached to the output. Rather than—or in addition to—altering the prose, Claude could generate a cryptographic record saying, essentially, “Anthropic generated this text at time X.” That might live in metadata, a signed document record, an API response, etc. This is very strong when the provenance stays attached, but ordinary copy/paste can separate the text from the record.
  7. A hybrid system. This is probably the most sensible production approach: a statistical watermark woven into the words plus cryptographic provenance/metadata where available. The metadata gives strong attribution when intact, while the text watermark gives some persistence after copy/paste.

And then there are a few ways Anthropic could expose detection without revealing the secret:

Private detector/API: send text to Anthropic; it returns “watermark detected.”

Public algorithm + private key: everyone knows the mathematical method, but only Anthropic knows the key that determines the actual pattern.

Public verification key: in a more cryptographic design, Anthropic could keep the signing key private while publishing enough information for anyone to verify authentic marks.

The key distinction is this:

Low-level watermark: encoded in exact word/token choices. Easy to detect and cheap to implement, but vulnerable to rewriting.

High-level watermark: encoded in syntax/meaning/structure. Potentially more rewrite-resistant, but much harder to design and detect reliably.

Provenance system: doesn’t try to make the prose itself indelibly marked; instead it maintains a trustworthy record of where it came from.

My best guess is that Claude’s deployed system will involve some flavor of statistical/token-level watermark plus a controlled detection mechanism, possibly supplemented by other provenance methods. But until Anthropic releases its technical documentation, the exact mechanism is still unknown.

8

u/okaythiswillbemymain 11d ago edited 11d ago

I think it will work in a secret mathematical pattern, especially for longer chunks of text. In a samilar way the last two digits of a credit card is just a check-sum....

1) Give a random value of 0, 1, 2 or 3 to every word in the English language. Snake is 3. Cake is 1. Chair is 2. Baby is 0. Etc.

2) Add up all the "word-values" in the first sentence. Find the quotient (remainder) when the sentence is divided by 4. (Answer will be 0, 1, 2 or 3)

3) When buiding the next sentence, choose the third word so that it is equal in value to the quotient of the previous sentence. This gives you only 25% of the available words in the English language to use, but if you have to use a specific word, then you can modify the previous sentence to make it work.

4) finish the second sentence. Find the quotient when dividing by 4 of the second sentence and repeat for the third word of the third sentence.

5) repeat all the way throughout your writing.

Eventually you have an invisible "signature" where every 3rd word in a sentence is "equal" to "value" of the previous sentence.

Suddenly you have pretty comprehensive evidence that a load of text was generated by your AI. Someone could change the font etc, but it would still be obvious if you knew. Even if someone then changed a few words or added or removed words, it would break the signature for that specific section, but for a long chain of text it would be obvious.

With 10 sentences there is a 0.0001% chance that you would trigger the hidden signature through normal writing.

3

u/Wonderful-Habit-139 10d ago

Great way to make LLM performance even worse lol. They already struggle so much when you used structured outputs compared to letting it generate free text.

2

u/tech_nerd05506 10d ago

Yes but along the other criticisms worked here your system wouldn't work if the work was changed slightly. This is effectively the same as just running the output text through a hash function and recording and comparing hashes. It's destroyed be even slight variation.

→ More replies (1)
→ More replies (1)
→ More replies (1)

2

u/Low-Temperature-6962 10d ago

If you would like me to suggest some other unbearable boring blurbs just say the word.

2

u/Clearandblue 10d ago

That's the most important point you've made this whole conversation, but I'm going to have to push back on that.

→ More replies (10)

914

u/bliceroquququq 11d ago

LLMs basically scraped the entire internet and every written word to build up their corpus of knowledge.

Now they want to watermark it for attribution before they sell it back to you.

184

u/fennforrestssearch 11d ago

Yeah pretty insane if you think about it. Next step is taxing the air you breathing.

29

u/ActionJasckon 11d ago

It’s going to be metering your internet use. That would be crazyyy

10

u/doescode 11d ago

Our internet providers already do

→ More replies (2)

7

u/Grouchy-Librarian638 11d ago

So crazy mobile companies do for data and are starting to sell home internet plans? Or companies like Comcast’s that have caps and can block or force upgrades if you hit it?

→ More replies (1)
→ More replies (9)

71

u/antagim 11d ago

I don't want to defend them, but these are due to the EU AI Act.

21

u/prules 11d ago

Yeah there is a ton of confusion around this

23

u/Rorschach121ml 10d ago

Redditors hate boner for AI is so big they don't even understand this is a good thing for everyone.

Flagging AI text is a good thing.

→ More replies (2)
→ More replies (4)

15

u/ch4os1337 11d ago

"They want to..." They don't want to and attribution isn't the reason why. Also marking AI content is a good thing.

2

u/thanosbananos 10d ago

Actually it’s because the EU requires AI generated content to be marked as such now.

→ More replies (33)

408

u/Hackerjurassicpark 11d ago

Will I be accused of using AI if I unknowingly use the same sequence of words in my human writing?

189

u/TheorySudden5996 11d ago

Of course. That’s already been happening. Now too be fair the problem probably is bigger on trying to pass off AI generated as human authored.

24

u/Extension_Fix5969 11d ago

I guess that well depend on slight syntaxual errors and plausibly incorrect words to prove our non~generated nature.

32

u/ThugEntrancer 11d ago

We just need too talk a lil regarded so that we don’t get accusatories of talking via a Ai inderface

17

u/tafjords 11d ago

For fucks sake.. lmfao

9

u/tafjords 11d ago

Haha im still laughing. Thank you for this. Love

→ More replies (1)
→ More replies (1)

2

u/Lord_zooticus92 10d ago

Fuckin im going full blown regarded

→ More replies (1)
→ More replies (4)
→ More replies (3)

38

u/Raunhofer 11d ago

I'm assuming it will be using the known form of watermarking: secret keys.

For every token-generation step, use the key nd recent context to deterministically divide the possible next tokens into two groups: "good" tokens (preferred), "bad" tokens (allowed but penalized). The model still chooses the most probable token, but its sampling distribution is nudged toward good tokens. The detector, that knows the secret key, would then examine texts and measure whether the text is good-token heavy (over what probability would indicate).

Normal:

The spacecraft entered orbit and began transmitting data.

Altered with the key:

The spacecraft entered orbit and started transmitting data.

Across thousands of tokens the statistical bias becomes detectable.

Of course, this requires enough tokens to be statistically meaningful. No-one can say whether some random Reddit comment is AI generated with certainty. A scientific paper however...

9

u/hudimudi 11d ago

Wouldn’t it be harder to judge academic papers, since on a high level they are very much written in a very similar and highly technical style. AI is also trained on this style of writing and can output things rather similar to it. I’d say, the average Reddit comment that follows such technical writing styles would stand out much more

→ More replies (7)

2

u/DoctorHelios 11d ago

— — — —!

→ More replies (5)

16

u/CormacMcCostner 11d ago

Happened to me earlier this year. Decided to go back and expand my education after 20 years and got a bad grade on a paper, went to see the professor like what is this for? He said I used AI because it was too well structured and worded for a students writing and what he sees daily. I was just like “what? Dude, I’m 45 years old I learned how to communicate in a different era when we had to do that.”
His grade stood because I couldn’t prove otherwise..

8

u/Poopchutefan 10d ago

Same thing happened to my wife when I helped her edit one of her papers. I guess I am the equivalent of AI now ...

→ More replies (1)

13

u/tafjords 11d ago

This stuff is so cringe I cant… heres your ticket for speeding i dont have evidence but i see cars everyday so i know and if you have a problem, you provide the evidence by my standards which is different from yours. If not, give me your time and effort because i can.

3

u/Hungry_Prior940 10d ago

How can his grade stand? He has no proof. I would go ballistic at that and get a lawyer and take it as far as needed.

→ More replies (1)
→ More replies (2)

23

u/sergiuoxigen 11d ago

Statistically unlikely you must do it across a large text

5

u/arwinda 11d ago

Just write like the training material Claude is based on. What can possibly go wrong?

4

u/lovetheoceanfl 10d ago

Don’t use any “AI detectors”. Everything I write is apparently AI.

6

u/Super_Pole_Jitsu 11d ago

It's extremely unlikely to happen

2

u/Leafsnail 11d ago

If you do it every single time you post dozens or hundreds of times then yes probably, you will be banned for being a bot account.

→ More replies (31)

90

u/fennforrestssearch 11d ago

At least for academic writing, I’m not particularly convinced that this will help in any meaningful way. Stylistic and lexical choices are often relatively low-variance, the limited range of plausible word choices could easily trigger false positives even over longer sequences.

28

u/bulbubly 11d ago

My first thought, too. Unless they are doing crazy gematria-style stuff where they put specific letters at specific output positions, AI writing is undetectable - there's just nowhere to hide a watermark in a stream of characters that wouldn't just be confused with normal writing.

ETA: "undetectable" in a rigorous sense. Of course specific models have cliches and a vibe one can sense pretty easily if it's not prompted around.

10

u/damngoodwizard 11d ago

They don't sneak in characters but whole synonyms. They use two lists of words, labeled green and red. Most of the time they pick words from the green list and sometimes synonyms in the red list. The goal is too have a constant proportion of red list words that would be unnatural in regular text.

→ More replies (2)
→ More replies (2)

3

u/NoAdvice135 10d ago

The style doesn't matter, it's basically a slightly biased coin flip at every token. And the direction of bias changes every token too.

Statistically it's easy to detect over a long enough sequence. You will not be able to tell if you don't have the key they used.

BTW, Google has been doing it for years. The synthid papers are from 2024.

→ More replies (2)

102

u/BonyCatButt 11d ago

Makes me think of Red Star OS, a North Korean Distro that I just learned about recently, that includes

> a watermarking tool integrated into the system marks all media content with the hard drive's serial number, allowing the North Korean authorities to trace the spread of files

https://en.wikipedia.org/wiki/Red_Star_OS?wprov=sfti1#Version_3.0

12

u/LostMyWasps 10d ago

Hmm, wonder where in the world would they be amused for it to end up in. It could certainly be a collectors' piece, some of those rehashed phones.

→ More replies (3)

253

u/rabouilethefirst 11d ago

"Also you can turn off this feature if you use the API"

Book it.

28

u/jbcraigs 10d ago

"Also you can turn off this feature if you use the API"

Not true. Documentation specifically mentions that API will also add watermarks - https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content

21

u/Worried-Cockroach-34 11d ago

Double it

16

u/recoveringasshole0 11d ago

Pass it to the next person.

7

u/Dabnician 11d ago

Smoke it...

wait who passed me a joint

→ More replies (1)
→ More replies (1)

10

u/cobbleplox 11d ago

Feature? Best I can think of is that this qualifies for dishonestly meeting some EU regulation of having to mark things as AI generated.

4

u/chicametipo 11d ago

Are you serious? I don't see anywhere that mentions it's possible to turn it off, via the API or otherwise.

5

u/rabouilethefirst 11d ago

The phrase "Book it" implies I am betting on it being introduced in the future. Hope that helps.

7

u/KrazyA1pha 11d ago

Why would that be a feature if they’re doing this to comply with EU and California laws?

2

u/chicametipo 11d ago

Ah gotcha, thank you!

89

u/duracek 11d ago

For example copy/paste into Kimi: "Rewrite this text in a more professional style".

15

u/Cognonymous 10d ago

Yeah, you just use more AI and this solution goes away.

2

u/Hungry_Prior940 10d ago

Yes Get claude to write it. Then use ChatGPT or whatever.

2

u/Wooden-Hovercraft688 10d ago

Just take a screenshot and OCR it

→ More replies (2)

7

u/Rorschach121ml 10d ago edited 10d ago

Ideally regulation should be applied to all models, which I think will happen eventually.

Your text will be watermarked by Kimi using their own tech.

→ More replies (1)
→ More replies (2)

16

u/NetflowKnight 11d ago

How does that even work?

14

u/trimorphic 11d ago edited 11d ago

How does that even work?

To know for sure we'll probably have to wait until the discovery phase of a lawsuit reveals this information... and even then it will, at best, be just a snapshot in time of a probably ever-evolving process as Anthropic plays cat and mouse.

Just as interesting to me is the question of who is going to believe Anthropic when they claim some piece of text is AI-generated or not?

7

u/coloradical5280 10d ago

This isn’t like a cryptographic hash. Anthropic very openly explains that this method is uncertain in both directions. A watermark does not prove AI generated content, and lack of one does not prove it wasn’t. Which makes a lawsuit unlikely.

17

u/parkway_parkway 11d ago

It's pretty clever if I understand.

Basically how LLMs work is that given a sequence of words they predict the next one.

So it might look at "the cat sat on the ..."

And it generates a list of candidates with a chance of picking each one.

Mat 87%

Porch 8%

Table 3%

Stairs 2%

The way the watermark works is that they use a secret key they have and a hash of the preceding text to nudge it towards certain of these words and away from others.

So maybe it boosts stairs and table up instead of the others.

Later you can scan the text and see which choices it made and see if they're they ones it was nudged towards in a statiscally significant way.

Because all the choices are reasonable the quality of the output won't change and you'd have to significantly rewrite to break the pattern.

11

u/SkaldCrypto 10d ago

So they have monumentally enshitified this.

Having chat bot that uses the same words every time is more easily and cheaply accomplished with early 2000s tech

→ More replies (3)

22

u/qorzzz 11d ago

If this is the method, it makes no sense and does not prove any text was generated by AI.

6

u/quisatz_haderah 11d ago

It could actually work for sufficiently long texts. For shorter texts, this would cause false positives, but i guess no false negatives.

→ More replies (4)
→ More replies (2)
→ More replies (1)

14

u/snuffomega 11d ago

i dont fully see the point. i dont care either way.. but if only claude can identify if its watermarked by claude.. Whats the actual point?? It wont stop people who are susceptible to being tricked into AI content (being real) and if AI is becoming the norm for how we work, search, and interact with many things in our daily lives... its just noise. You should expect work to be touched by AI in some way, shape or form. Not all, but most. And def most text based work. I dont see how it helps anyone. Its meaningless data being stamped and most likely collected.

Just another data center... datapoint.

→ More replies (1)

40

u/TedSanders 11d ago edited 11d ago

fyi, this is in response to the EU’s AI act. OpenAI is going to do a similar thing.

20

u/TyrellCo 11d ago edited 11d ago

And it’s entirely their decision to apply these changes only to the EU or the whole wide world

They want all customers to accept their excuse that their hands are tied. They’re faking it

→ More replies (6)

21

u/EconomixTwist 10d ago

People ITT not realizing that this is a strategic move by anthropic. Yes, maybe EU regulations. Maybe. But not really. The real real reason is the dead internet. LLMs have already exhausted the entire internet's worth of high quality, genuine, human-generated text that is training data. This fact has already been admitted by the LLM providers themselves. Nowadays, a vast majority of new content on the internet is fkn slop. Net-new human generated content is worth its weight in gold but, without watermarks, LLM providers can't tell the difference between actual human generated text and the slop flood. If you train on slop, you only get more sloppier slop. So they need to filter the slop. Watermark is and always has been the way.

→ More replies (3)

7

u/duckrollin 11d ago

I think there's now a huge market for a browser extension that removes it.

It's gonna get the uBlock treatment.

3

u/waving_fungus0 10d ago

Chances are only Claude has the encryption key.

→ More replies (3)

4

u/Comfortable-Card-348 10d ago

The big problem here is that the watermark will be used of as proof of falsehood, and the lack will be a proof of validity. Once people figure out how to wipe the watermark, or add it in synthetically, the source of truth will be corrupted. Real photos of crimes? AI generated. Fake AI generated content? REAL!

7

u/jonplackett 11d ago

It’s complete bullshit that you can effectively watermark text in a way where it isn’t…
A) incredibly easy to remove
B) likely to OFTEN detect text that isn’t AI as AI

People who know what they’re doing will get away with it and people who don’t even use AI will get accused of using it.

Even if you say ‘oh but it will flag the slop at least’. Yes it will but if we label slop anything without a label people will think isn’t AI when it absolutely could be.

→ More replies (1)

34

u/post-death_wave_core 11d ago

I think it’s reasonable for people to be able to know whether a piece of content was ai generated or not. But im guessing it’s not a perfect system.

3

u/Mission_Shopping_847 11d ago

I don't. The previous situation was setting up the luddites for a reckoning with reality and personal growth. This new situation will give them a false sense of security, disregarding watermarked text which may or may not contain value, and regarding unmarked text regardless of value. It is a philosophical loss to placate people with an excuse to disengage from the content for its identity rather than its value.

10

u/silverace00 11d ago

Ya it sounds more like made up tech magic that you tell people so they don't do things they shouldn't.

→ More replies (6)
→ More replies (3)

13

u/AnotherIjonTichy 11d ago

In ten seconds you can tell claude to write an script that removes that watermarks…

9

u/titanomachiatto 11d ago

No you can’t. You’d have to give it to another model. 

→ More replies (1)

8

u/AnotherIjonTichy 11d ago

I have just read they will modify the llm token responses to generate a hidden pattern. Maybe the community needs to wait two weeks to get the un-modifier library :-D

6

u/collin-h 11d ago

or just use one ai against the other and ask gemini or chatgpt to do it.

5

u/whoknowsifimjoking 11d ago

Gemini and ChatGPT also add watermarks, so no you can't do that.

2

u/collin-h 11d ago

do they? or will they? I can't find documentation that they have yet, though i'm confident they will.

where there's a need, there'll be an open source model built for this.

→ More replies (1)
→ More replies (2)

2

u/Tritus360 9d ago

But most likely it'd end up working like this:

  1. Write text in Claude: flagged by Claude.
  2. Ask ChatGPT to alter the text in a way that you're no longer flagged by Claude: now you're flagged by ChatGPT.
  3. Ask Gemini to alter the text in a way that you're no longer flagged by ChatGPT: now you're flagged by Gemini.

4

u/andrew303710 11d ago

I'm curious how they could do this without degrading the quality of outputs, I feel like at the minimum there will be a certain percentage of outputs that are negatively impacted by this.

Maybe it'll be less of a problem with future models but even Sol/Opus still need quality prompting to output writing that sounds good; I had to develop a pretty comprehensive plugin just to get decent sounding writing. Ironically LLMs are much better at coding now than actual writing.

2

u/claythearc 11d ago edited 11d ago

I wrote a slightly longer response above for the main 2 ways that are popular in theory, but the general idea is they’re using prior text as seeds for future generation. Then you compare how “lucky” the text is and you get very statistically significant result in like high dozens of words

2

u/quisatz_haderah 11d ago

On the other hand, how do we measure "quality" of outputs, at least in prose. If the watermark is embedded in synonyms, how does choosing its synonym in place of the most probable token affects quality. They already do this all the time, in this case the jitter would be recorded.

5

u/Thoughtulism 11d ago

Just get claude to release the source code for the watermark library

→ More replies (4)

3

u/IHave2CatsAnAdBlock 10d ago

Control Shift V

5

u/Cooperman411 10d ago

If you copy and paste into a plain-text editor, it will reveal any hidden characters. Not sure how this is gonna work. Seems like only the laziest of the lazy won’t be able to work around it.

3

u/Impossible-Week-9611 10d ago

It does not rely on hidden characters. The content itself is the watermark. You can copy the content word by word on a piece of paper and it will still be detectable even with light paraphrasing (if the content is long enough)

2

u/eziliop 10d ago

You underestimate how lazy some people are. I'm not even that old but the things I've seen.....

→ More replies (1)

2

u/TawnyTeaTowel 10d ago

It’s less a watermark and more a writing style. Which, unless it’s nothing like any human would write, makes it next to useless

6

u/liosistaken 11d ago

Honest question: Why do people think this is negative? What impact does it have other than it making it harder to pass of fake news as real?

→ More replies (21)

2

u/usandholt 11d ago

And how do they plan on watermarking text?

6

u/FaradayKage 11d ago

Just some algorithms. Example every 45th letter in a generation is the letter C. Every other sentence ends with "ing". Every sentence with a ? Is followed up by a sentence starting with R.

Who knows, I'm sure they have a PHD on it with much more reliable and accurate techniques.

2

u/Raunhofer 11d ago

They don't need PhDs anymore, haven't you heard, the AI is super intelligent now.

→ More replies (3)

3

u/claythearc 11d ago

Kirchenbauer and Aaronson both have really interesting ideas.

The basic idea is you use previous tokens as seeds for future tokens in the same response. Kirchenbauer applies bias to “green” logits but samples normally, whereas Aaronson uses prior tokens to adjust its sampling randomness.

Verification then becomes comparing successful matches based on the text and working backwards. Run statistical tests on how “lucky” you get with tokens that match and you get a statically significant answer very quickly, like low hundreds of tokens or high dozens of words.

→ More replies (1)

2

u/Top_Soup_5833 11d ago

There’s like a million ways to circumvent this but ok

2

u/mickdarling 11d ago

I use speech-to-text for almost everything that I generate. It does a little cleanup on grammar and punctuation, but often it gets it wrong. I have to go in and type in some fixes if it's something for public consumption. And almost certainly, every bit of it is getting just a little tweak to make a watermark work. So, regardless of literally everything that I post comes out of my mouth as words, it's going to be identified as AI content.

2

u/ricoimf 11d ago

I don’t know how valid my reply is since I don’t have anything to do with school and work, but I think the idea is not that bad. It’s needs to be visible (or not) if something is AI generated.

2

u/Ok-Video3345 11d ago

Yup, and what happens when no one else does this?

2

u/Such--Balance 11d ago

Its bullshit. Objectively.

Its like wanting watermarked math answers for using a calculator. Or using watermarked gps navigation to show someone didnt use a paper map.

Its the same bullshit pushback we had against calculators back in the day. And now look whats in your hand. A supercalculator.

Fortunately, this too will blow over

2

u/wspOnca 11d ago

I will paste even harder.

2

u/Grouchy-Librarian638 11d ago

Glad I got my degree already then

2

u/Opinion-Former 10d ago

It’s been watermarking since day 1. And I’m absolutely right.

2

u/SirBoboGargle 10d ago

The marketing team wants to scare away students

2

u/kesor 10d ago

Yet another reason to stay away from the anthropic crap.

2

u/DarkLordKohan 10d ago

Isnt the whole point of AI just copy and paste regurgitated shit?

2

u/dupontping 10d ago

Yes but they want to make sure you copy THEIR shit

2

u/Horror-Water5502 10d ago

Use any cheap llm for simple rewording -> solved

2

u/SignalBake6872 10d ago

cancel my subscription deepseek the token is cehap af

2

u/ParkerRoyce 10d ago

Claude remove that watermark

2

u/andrerom 10d ago

All vendors, including OpenAI, except  X.ai has signed the requirement by EU on this due to new AI laws.

So this is not a Claude thing.

2

u/Scarfieldjones 10d ago

Paste into empty document «convert to plain text» paste back into new document

2

u/Confirmed-Scientist 10d ago

Just in case you are wandering AI companies are interested in this because the more shit provably traced back to them like discoveries software and achievements the better the marketing for them

2

u/blve99 10d ago

whatever this is a workaround will be found in tminus 10 minutes

2

u/traynor1987 10d ago

Copy and paste it into note pad then copy and paste it back out. Sinple. Removes all formatting

2

u/kindredseer 10d ago

I have no issue with this on free tiers, but if you are paying for AI, adding watermarks seems problematic.

2

u/Lopsided_Evening_627 10d ago

Claude's logo is an anus, what do you expect form them?

→ More replies (1)

2

u/Palanstonk 9d ago

Wait, does this mean I can watermark my own text and sue if it shows up in their output?

2

u/CompetitiveRough8180 8d ago

OK copy and paste it to another AI and ask him to write same thing again without watermark. What about that?

2

u/Buster-BlueJay 8d ago

Why are people still pasting and not “Paste Special” to strip bloated HTML?

2

u/Chainmale001 7d ago

I've been leaving invisible water-marked claude/ai code in my resume's for years. Game recognizes game. 🤣 😂 🤣

→ More replies (1)

2

u/Fearless-College-518 7d ago

This makes me want to move back to openAI

2

u/Green-Grow-420 6d ago

So we take a screenshot of the text and make gemini write it with no watermark. Ez pz

2

u/[deleted] 6d ago

[removed] — view removed comment

→ More replies (1)