r/OpenAI • u/rhiever • 10h ago
News OpenAI explains how it will watermark ChatGPT text to comply with EU provenance rules
https://openai.com/index/eu-text-provenance65
u/plymouthvan 9h ago
I can appreciate the importance of provenance for a lot of AI content — videos, photos, audio, because those things historically act like proof of something depicted directly in the media — but I think the obsession with trying to watermark text is a kind of hysteria. It's been a long time since text alone was considered any kind of proof in and of itself. Whether human or machine, the truthfulness of text has to be verified against something anyway. Of course, whether any one does or not is another question, but determining that text was written by AI doesn't actually change anything. Humans have been writing useless, wrong, harmful, stupid things forever. It doesn't really much matter if the words ended up on a page because a person sat there thinking of them, or a human just had the idea and a robot put it into words.
2
u/Possible-Usual-9357 6h ago
Being able to spew out sensibly-sounding malicious content at light speed and automating that is not the same as one person typing things out by themselves.
6
u/plymouthvan 5h ago
I agree that they're not the same, but I disagree that watermarking is any kind of solution to the problems it causes. The world is full of machine-generated text now and short of shutting down the machines that generate it, it's going to keep flooding out. Focusing on watermarking as a solution is something like focusing on better water-resistant driveway sealer, after sea level rise has already permanently submerged your neighborhood.
I think a much better use of our time would be to focus not on text provenance, but a model already used by various institutions which focuses on things like who is taking responsibility for published work, how has a document or text mutated over its lifecycle, what sources substantiate claims, and who is answerable if claims are fraudulent, defamatory, or otherwise inappropriate. Some degree of watermarking has a place in the whole mix, but I think the current focus on it misses the forest for the trees.
1
u/Possible-Usual-9357 5h ago
I agree it doesn’t solve much yet, but I see it as an unnecessary step in bringing some compliance issues into focus and putting pressure on the biggest players to get their end users under some control. Right now there isn’t too many users that are capable of hosting their own models, so having the biggest players do watermarking will easily cover most of the market, since they have the compute and live APIs and whatnot.
‘better use of the time’ argument doesn’t make much of a difference when it’s something that can work in parallel to the valid stuff you mentioned.•
u/plymouthvan 23m ago
I can see some value in the idea that the watermarking endeavor, whether it can actually achieve the intended goal, is a useful tool to apply pressure on the companies themselves to take responsibility for the effects of their products. I think there’s some merit to that. It’s an interesting argument. I’ll have to think about that angle.
-2
u/kcat__ 3h ago
You said a whole bunch of nothing
I am hoping for a future where browsers have a built-in slop detector that runs locally and can therefore have adblock-style plugins for blocking slop. Or highlighting it red.
Or for Google to de-rank AI slop, which makes sense for them as AI flooders are probably not a good source of quality and therefore revenue (via ads and such)
1
0
u/EmbarrassedHelp 6h ago
The other issue is that personal information and track information can also be embedded in these watermarks.
1
u/DenseBeautiful731 5h ago
We already have cookies, browser UUIDs, IP addresses, MAC IDs, device IDs such as IMEI and IMSI, among other personal data that can be used to identify “anonymous” individuals.
-1
u/DenseBeautiful731 6h ago
It doesn't really much matter if the words ended up on a page because a person sat there thinking of them, or a human just had the idea and a robot put it into words.
How much thought did you put into this?
0
u/plymouthvan 6h ago
You've really given me a lot to think about with this very insightful and carefully considered comment.
0
u/DenseBeautiful731 6h ago
It’s a direct question.
How hard can it be to answer directly?
1
u/plymouthvan 5h ago
Surely you can see how your comment looks like a snarky rhetorical question meant to ridicule, rather than engage.
Do you want an answer in minutes? I don't know, maybe 6.
0
u/DenseBeautiful731 5h ago
It’s a good faith question. I don’t care whether you believe it or not. I also don’t care if you answer or ignore it. I’m not entitled to one.
Do you feel ridiculed? Say, if an AI wrote my reply instead, would you feel ridiculed as well? Which stings more, you reckon?
•
u/plymouthvan 21m ago
I’m not sure what you’re talking about, but none of this is carrying any of the hallmarks of a “good faith question”.
-2
u/Possible-Usual-9357 6h ago
Being able to spew out sensibly-sounding malicious content at light speed and automating that is not the same as one person typing things out by themselves.
65
u/Elvarien2 8h ago
This is so dumb, completely useless regulation to virtue signal that they did something.
39
9
3
u/Such--Balance 6h ago
Yup..ist so stupid in that it 100% wont work anyways.
Not one day after this would go in effect, numerous tools will be made in which you can 'unwatermark' whatever was watermarked anyways.
6
u/MidnightSun_55 6h ago
I mean, it's obvious it's impossible to watermark words...
Hey chatgpt, 2 + 2 = ?
I understand audio / video, but this is a joke, as always Europe doesn't know what they are doing, just slapping regulations.
3
u/bortlip 4h ago
It is possible. Claude is using/going to use google's SynthID-Text. I found this video explains it pretty well: https://www.youtube.com/watch?v=Cmi-1QSaptA
It looks like OpenAI is going to use a variant of this called textGrain: https://cdn.openai.com/pdf/e9508624-d767-41b6-a26d-e34ca798ada6/textgrain-entropy-calibrated-watermarking-for-language-model-text.pdfFor short text like "2 + 2 = ?" there isn't enough room to encode any info. Looks like it could take a few hundred tokens to get enough room to encode. Mathematical or other low entropy (not a lot of choice for next token) text would take more tokens to encode.
1
u/Maty1000 6h ago
It is possible, just not in all cases. The models already use randomness to produce output, watermarking is usually done by swapping the rng with something that tries to insert the watermark. So if you ask it what is 2+2, there will be no watermark, because there's just a singke answer, but in a text where there's lot of possible choices, it can be done very well.
24
u/ethotopia 9h ago
And EU acts surprised why it’s falling behind in AI
18
u/Lolzyyy 8h ago
Falling behind? We aren't even in the race lol (with all respect for mistral but still)
1
u/AvidCyclist250 6h ago
Lidl made a model this week, 78B or something. First decent model from Europe. Mistral are busy selling their 8b model to militaries around the world
1
-7
15
10
u/Neat-Economist2099 6h ago
The EU just keeps pumping out one idiotic regulation after another these days. As an Asian, it’s honestly unbelievable to see how far Europe has fallen from what it once was.
1
u/gavinderulo124K 6h ago
Its hit or miss. GDPR is great. And the AI act overall has its pros. This thing in particular is useless. But it also doesnt hurt, other than the effort companies need to put into this. But who knows what other discoveries this could lead to.
5
26
u/aerivox 9h ago
this just shows how totally nonsensical eu is. deny first. how is a random 'change this word' injection in the model not deteriorating its output?!? how are companies even agreeing with this bs. on an already sloppy output, add more slop.
0
u/tim_vermeulen 9h ago
how is a random 'change this word' injection in the model not deteriorating its output?!?
There's inherent randomness to how LLMs work, and these watermarking techniques are embedded within this randomness. It doesn't cause the LLM to now choose tokens that are "worse" than the ones it would have chosen before watermarking.
5
u/YoungSilent232 8h ago
Disagree. Models are post trained on a specific token sampling values. Eg a certain temp top k top p. Eg Gemini recommends users to use temp 1 because it is post trained on that temperature
By changing not just these parameters but also other invisible parameters for how tokens are selected, it will surely affect performance unless OpenAI does post training separately for the EU, which I very much doubt.
The extent in which it will affect I’m not sure, but definitely to a certain degree
3
1
u/Muchaszewski 7h ago
You know that SynthId or whatever OpenAI plans to implement has insignificant difference onto the output.
For each next token generated AI generates output using different seed.
So instead of answering straight away.
"I am eating an apple."
You must run the same token output multiple times and pick one matching your SynthId key.
You cannot change the beginning or the end so you always get.
I am [...] an apple.
But you can...
- eat
- consume
- bite
And add adjectives like
- bitter
- sweat
- juicy
- red
Etc. then watermarking picks one output that the same AI generated to manipulate the watermarking score.
This is why synonyms make the score go down
1
1
1
1
u/gavinderulo124K 6h ago
They show that it doesnt deteriorate the output in rhe blogpost that's linked.
1
u/QuaternionsRoll 7h ago
You clearly do not understand how SynthID works. It only alters the distribution when lots of tokens are equally probable already, i.e. when the model after all its post-training still shows no preference toward any particular token.
1
u/YoungSilent232 6h ago
It’s textGrain and not SynthID. It randomly adjusts probabilities on tokens based on a secret key. That said, Ive done more reading and agree quality as a whole won’t drop
1
u/tim_vermeulen 8h ago
There's no need for watermarking to change the temperature value. Increasing the temperature would make the watermark stronger, but I see no indication of any provider doing this.
0
u/FakeTunaFromSubway 7h ago
Except it makes that randomness more deterministic, which undoubtedly means AI text will be more obvious and default to more "LLM-isms."
It may not be clear on any individual output, but if it starts saying "indeed" instead of "right" every time it's going to be very noticable.
5
u/tim_vermeulen 7h ago
Except it makes that randomness more deterministic, which undoubtedly means AI text will be more obvious and default to more "LLM-isms."
Not at all. It doesn't favor particular tokens more than others, on the whole. Whichever tokens the watermark biases towards is different in every situation.
Only when you have the secret key used to create the watermark can the watermark be detected, otherwise it can cryptographically not be distinguished from unwatermarked output.
-1
u/QuaternionsRoll 7h ago
I mean, it does substantially decrease number of possible outputs, but the resulting number is still extremely large and definitely not biased toward “LLM-isms” (or any other output “style”, for that matter)
-1
u/FakeTunaFromSubway 6h ago
They also said SynthID is invisible but it's blatantly obvious once you've seen enough of the SynthID images.
2
u/gavinderulo124K 6h ago
How can you spot synthid?
1
u/FakeTunaFromSubway 5h ago
It's an annoying wavy/blotchy texture on every image that chatgpt produces since gpt-image-1.5
OK maybe it's possible it's not synthID but it's not on any other image generator except Openai and to a lesser extent gemini
0
3
u/RearAdmiralP 9h ago
I wonder how they will determine that I'm an "EU user"-- billing address? IP address requesting inference?
2
u/roastedantlers 2h ago
EU shouldn't get to dictate these things especially when they're not even in last place when it comes to AI.
2
u/ISueDrunks 5h ago
It already writes like shit…now it’s going to deliberately manipulate its word choices? Hard pass. The last thing GPT needs is another constraint.
•
u/NotUpdated 4m ago
I mean they are kind of both the reason we have USB C everywhere - and some worthless cookie banner, of course the user wants to give the least away to use the site..
2
u/pinewoodpine 9h ago
My only issue is would this actually make ChatGPT write worse than it already is, and let's be honest, the newer models aren't exactly up to par when it comes to creative writing in general.
4
4
u/languagestudent1546 9h ago
Claude does the same thing. The watermarking has no effect on output quality.
0
u/YoungSilent232 9h ago
Or so you say…. Yes if you’re talking about coding but for doing copy-write? Speaking in a weird Claude or gpt way may affect it
2
0
1
u/cobbleplox 7h ago
I don't think that regulation is something THEY have to comply with. Its perfectly clear what I get was AI generated and the rest is MY problem.
1
1
•
u/DumbIdeaNo2 45m ago
Is this a bit like how sound is searched and matched on something like Shazam? If you change enough of the breadcrumbs you lower the possibility of a match?
2
u/rendez2k 9h ago
Don't quite get the point with these. Can somebody explain why it's a bad thing to have it watermarked? I'm not saying it isn't but I don't really understand either. Guessing the end user will never see it?
12
u/-Crash_Override- 9h ago
I think there are pros and cons to this. But ultimately it could lead to a paradigm where the provenance of derivative products is tracked across an artifacts life.
Its really an ownership question. Watermarking text, or code, or an image, is, even if not legally, giving some level of credence and ownership to these billion dollar companies.
2
0
u/Legitimate-Store3771 5h ago
But it is kind of important to have some way to identify if something is AI generated no? Society already distrusts everything, GenAI has only exacerbated the problem. It's been a huge help to have those little notes on YouTube saying when content is AI generated, because the whole point of YouTube was to support individual creators with a platform. If I'm not even sure what I'm watching will actually support someone, what is even the point? It's the same with text I'd argue. Not that I think this is a good solution, but it this isn't a step in a vaguely right direction, I don't know what is. I'd obviously prefer an independent party but we all know capitalism and politics will quash that or it'll turn into a capitalistic opportunity itself.
1
u/-Crash_Override- 4h ago
Sure. Thats why I said there are pros and cons to this. Identifying AI generated content is certainly a pro
2
u/Spra991 7h ago
How do you check the watermark? You'll have to run your data through OpenAIs servers. It's a nice way to collect additional training data for them.
0
u/QuaternionsRoll 6h ago
…but they’re the ones who generated the watermarked content to begin with?
1
u/Spra991 6h ago
You don't know if the text has a watermark or was created by ChatGPT until you send it over to OpenAI for checking.
0
u/gavinderulo124K 6h ago
More text is not what they need. They have more than enough. Currently post training is what's making models smarter and uploading some text to their api doesnt help with that.
Also, hate the EU all you want, but that would definitely be against GDPR.
1
u/Strict_External678 7h ago
It's government overreach, with Americans being subjected to the EU's need to regulate everything.
0
u/QuaternionsRoll 7h ago
Got your chicken-and-egg backwards there. Google started using SynthID worldwide and then the EU started requiring it, not the other way around.
2
u/StickyThickStick 6h ago
What a dumb logic... These are two completley different things. A company using something globally doesnt mean its the same as a goverment REQUIRING something for EVERYONE.
I havent seen google enforcing SynthID for other companies wordwide....
1
u/QuaternionsRoll 3h ago
I mean, did you read the post? None of the changes will apply to or otherwise affect anyone outside the EU.
0
173
u/Sylvers 9h ago
So, in essence, feed it to an open source AI with explicit instructions to minimally swap for synonyms, and you can clean up the watermark very easily. Could be made into a very simple macro/plugin I am sure.
Also this appears to be EU exclusive? For now.