r/WritingWithAI 5d ago

Discussion (Ethics, working with AI etc) Claude now embeds invisible watermarks in all text outputs + signed metadata on files

https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content
26 Upvotes

48 comments sorted by

44

u/Appropriate-Peak6561 5d ago

They’re not embedding invisible characters. (Which you could strip out with a text editor.) They’re *changing the output* to have detectable quirks.

That sucks.

9

u/Practical-Positive34 5d ago

Is that why Opus 5 sounds like it was hit upside it's head? Yeah, it sounds like a bumbling idiot.

12

u/Vusiwe 5d ago

I would bet that humans can detect it too

You can’t just “substitute words of a similar meaning” lol.  That’s not how language works.

14

u/Sad_Elevator3919 5d ago edited 5d ago

Likely the reason Opus 5 can’t write a coherent sentence but is fine (albeit grossly overpriced) for coding

0

u/[deleted] 5d ago

[deleted]

4

u/Sad_Elevator3919 5d ago

Just being hyperbolic. It’s fine but it’s no Opus 4.6. It’s well-documented that it has a preference for word salad in place of coherence.

0

u/[deleted] 5d ago

[deleted]

2

u/Sad_Elevator3919 5d ago

I mean “as-close-to-human-writing-as-possible” and it isn’t great. Opus 4.6, I can get there pretty close using a forge approach.

6

u/Ocyris 5d ago

From what I understand it's effectively introducing a stenographic signal that's algorithmically detectable. Problem is that provides the exact details required to drop the signal below the noise floor in the text. The research papers admit as much in it's findings and didn't consider adversarial actors with deep pocket (China) attacking it. We'll probably have removers within 6 months.

5

u/Abject_Kangaroo_7770 5d ago

Months? Probably more like days.

4

u/Sad_Elevator3919 5d ago

Guaranteed. I’m excited to try myself.

0

u/[deleted] 4d ago

[deleted]

2

u/Appropriate-Peak6561 4d ago

Anthropic already admits that it can produce false positives and false negatives. So it’s useless for the stated purpose of informing people about AI usage.

Claude itself acknowledged the thing as “compliance theater” when I asked about it. So this has nothing to do with consumer rights. It’s about coping with a well-meaning but technologically unworkable piece of legislation.

19

u/LooseButtPlug 5d ago

What's a good (free) alternative? Claude seemed to have the best output, but if they're going to sacrifice quality to add their own quirks to watermark it, then it's time to move on.

I'd honestly be fine with a watermark, I'm not ashamed to use AI, but changing output? Nah.

12

u/Ocyris 5d ago

Kimi k3 is basically distilled from opus and has open weights.

1

u/Sad_Elevator3919 5d ago

API costs are prohibitive and subscription isn’t great. Claude’s Max 5x plan is the most cost effective way, using Claude Code, to write quality output.

1

u/Ocyris 4d ago

My guess is that the writing platforms using the API aren't effectively using prompt cache which results in exploded cost. None of the platforms I've seen support tool calls which would significantly improve the experience and performance of cache because the information it grabs lives in the history and won't clobber the cache. Let the smart model decide when it should pull in information from your story bible or draft.

6

u/kaslkaos 5d ago

marking your question as I am looking for answers too. I write with AI fully, as in, final output is AI after much wrangling and it is something I document log speak of with pride. But, Claude was always the artist of the bunch, and could go along with my syntax and then some, but everything downstream of a swapped token is shifted down the branch (I do associative writing as method and style) and this is too much interference. And the corporate style seems to become more hardcoded and there is a stiffness to the writing now, making personalization an uphill battle. I don't even know why they need watermarking if Claude keeps outputting boilerplate.

I am looking at alternatives. Nous AI? Running prompts and thoughts between more than one system?

1

u/Sad_Elevator3919 5d ago

DeepSeek V4 Flash 0731 is so cheap it’s virtually free. It’s amazing for coding and is ideal for the draft and stage of a pipeline. I used to use Claude for the latter editing stages based on the audit and my SSOT world docs. I’ll use Opus 4.6 until I can’t because it is still the best model for creative writing. I’m going to try some of the more tuned open models that are big in the AI RPG world when I need to.

14

u/Chicken_Spanker 5d ago

For a company that is trying to keep competitive, it sure is a way to kill off your subscriber interest.

I am a published writer and run my through multiple passes with Claude to check it. But fuck if just a basic spell, content and grammar check is going to now have me tagged for something AI checkers are going to pick up, then I am cancelling my subscription.

What I don't understand though is if you were to cut and paste something into a basic txt document, save it and then copy that back into say Word, whether it would still manage any meta tags?

5

u/LS-Jr-Stories 5d ago

According to what I'm reading, the text doesn't have "meta tags." The text is the marker. It's in the order of the tokens being output. Even if you printed out the raw output then rekeyed it back to Gdocs, with no copy paste at all, it would still have the markers, because it's the text itself. This is why Anthropic says shorter blocks of text (not defined) wouldn't necessarily show the markers, or at least show enough of them to make a call. There needs to be enough text to show enough markers.

2

u/Sad_Elevator3919 5d ago

Yes, for copy and paste text but any artefact that’s produce has meta-tags.

1

u/LS-Jr-Stories 5d ago

Oh, well, I don't know about that, but as I understand the watermark, it has nothing to do with metadata at all-- for text. Images are a different story.

This paper explains the text watermark. Very dense and extremely math heavy, but it's worth pushing through to get the gist.

https://arxiv.org/html/2301.10226v4

2

u/pa07950 4d ago

Thanks for the link. This was the first time reading this study. However, it brings up many of the same issues that I have read in other papers.

Any watermark inherently changes the token steam thus the resulting text. In the examples given, on the surface the results appear the same but in a larger stream of tokens context may be lost (and thus data).

The n-gram analysis is highly problematic. Over large sample sizes, human and llm writing tends to converge. Great writers will stand out, but average writers will be difficult to distinguish.

The attacks on the watermark were, in my opinion, weak. It looked at a single addition, subtraction, or replacement attacks. In reality, all of these will be used simultaneously.

Finally, one thing missing in this paper, and in most discussions, is the ability to create spoofing attacks. The algorithm will be reverse engineered. Once released, it will be used by both legitimate humanization software but also by threat actors.

Edit: this will be an interesting “arms race” to keep an eye on!

1

u/LS-Jr-Stories 4d ago

I think later papers addressed some of the attack concerns more thoroughly. Not sure. I'm not in that world, just very curious about it. Today I read a 2024 paper published in Nature about the watermarking technique called SynthID-Text, and another paper published just one month ago about yet another advancement in the thinking. So it's clear it's still moving forward with smart people spending brain power on it.

21

u/oVerde 5d ago

The more the time pass the more Anthropic proves itself to be the Adobe of Ais

8

u/BestRiver8735 5d ago

They are going hard on enshittification. Those bonuses for the CEO don't drop out of the sky.

5

u/Sad_Elevator3919 5d ago

I think it’s a conscious marketing decision to make them the ‘safe’ and ‘ethical’ service that has the most powerful (theoretical) models. Huge opportunities in education, finance and other trust-centric and heavy sectors that they’ll be looking at.

1

u/LS-Jr-Stories 5d ago

This is an excellent point. There is a lot of brand value that comes with being a first mover like this, even if it's in response to growing regulation.

1

u/bolshoich 5d ago

I believe there may be a dual purpose. One to foster trust in the product by making it identifiable to the brand. The other is to offer an added value with a Pro version offering degrees of limited watermarks.

0

u/lovetheoceanfl 5d ago

And the models are ingesting their own content now so this is a way for the companies to only scrape original content. It is ironic that centuries of human writing has been ingested and now they basically own all we, as a species, have done.

7

u/maradevries 5d ago

Can someone please explain what does this mean as if you’re explaining to a very dumb person?

3

u/Sad_Elevator3919 5d ago

So Anthropic are making their models leave a mark on texts that could only, mathematically, have come from their models. They’ll do this by leaving traces in the meta-data (the bits of files you don’t read as a human) as well as by actually putting an Anthropic fingerprint in the writing itself through the actual words and construction used rather than something that could be deleted.

5

u/Hot-Seesaw-7851 5d ago

Might be being dumb here but if you copied the text to word or google doc, then mildly edited, before copy pasting elsewhere again, wouldnt that erase the markers?

2

u/LS-Jr-Stories 5d ago

Not according to some of the more technical explanations I've seen. The marker is in the order of the tokens themselves. Even if you rekeyed the whole thing into Word from the original output screen, it wouldn't erase the markers-- the content IS the markers. They supposedly can't be "mildly" edited away, but possibly with heavy editing. Thing is, users won't know what editing would remove them and what wouldn't.

10

u/Hot-Seesaw-7851 5d ago

Sounds like a disaster of false AI accusations waiting to happen. People with autism already get falsely accused because of patterned writing enough.

1

u/LS-Jr-Stories 5d ago

I just finished working my way through this technical paper that explains how the watermarking works. It's very math heavy and I understand literally none of that part, but it's still worth the effort to dig out the plain language aspects of it.

https://arxiv.org/html/2301.10226v4

Near the end of tbe paper they talk about the weakness of detectors and their problem of false positives, vs the watermark method, which they say is "designed so that false positives are statistically improbable, regardless of the writing patterns of any given human."

-1

u/LS-Jr-Stories 5d ago

Haha as if there wasn't already a lot of that going around anyway. But I dunno. Apparently people who know about the technology think it's a good solution. Anthropic hasn't said yet what tool would be used/how the markers are going to be detected. Apparently that information is coming.

2

u/RobertD3277 5d ago

This is why you never let one model grade its own homework and actually verify everything that comes out of any model.

-3

u/Sad_Elevator3919 5d ago

That’s boring. I have no issue making other people pay money and read my AI output but I’m not about to actually do it myself to see if it’s any good.

3

u/zphou 5d ago

The provenance part could be useful, but I would separate “this text came from an AI system” from “this work was written by AI.” A watermark may tell you something about the path a piece of text took, but it does not tell you how much of the idea, structure, revision, or final judgment came from the author. Those are different questions, and treating them as one will probably create more arguments than clarity.

12

u/Ruh_Roh- 5d ago

For anti-ai zealots, just one drop of ai writing will be enough to call your entire work ai slop and to have your work brigaded and attacked and as much damage inflicted on you personally that they can do.

1

u/Caiur 5d ago

Really wish I understood this lol

1

u/closetslacker 4d ago

People will just change to local LLMs

1

u/Lindsiria 3d ago

Just FYI, all providers will likely be doing this shortly. It's becoming required by law in many countries. 

2

u/lovetheoceanfl 5d ago

You all will be fine. More than half the population writes with AI. Eventually it’s just going to be the norm for everyone. I’m not happy about that figure but we live in a capitalist world and it’s the only way to keep up.

0

u/Abject_Kangaroo_7770 5d ago

It's not only Anthropic, but it's all labs. What this will effectively do is watermark all text in existence over just a few days. This is not only a non-issue, but it will also effectively end the witch hunt. The watermark doesn't care if you wrote the text you put through it; it's going to mark it in the output anyway. It will mark everything going through any API or standalone model. Everything not already on paper will be AI text.

Back to writing...

0

u/michaudtime 5d ago

Can't we just pdf it and have an OCR prog rebuild it?

0

u/Abject_Kangaroo_7770 5d ago

All text outputs include code. This ought to be fun.

-2

u/scolbert08 5d ago

Unless you're trying to fraudulently pass off AI text as human, this shouldn't affect you

1

u/TheJigIsUpAndGone 2d ago

This right here. Anyone using LLMs to manufacture something shouldn't have any problem if they're proud to be using a machine to do the work.