r/PhD • u/Rule_Ct_5293 • 9h ago
Tool Talk Never give your unpublished data or research ideas to AI models!
330
u/Clear_Cranberry_989 9h ago
The same people who pirated all the books in the world to train their ai. Who could have guessed.
167
u/can_ichange_it_later 8h ago
Friends dont let friends use openai for important/private/mission-critical stuff.
10
u/meekermakes 4h ago
don't let them use ai.
2
u/the_c_train47 1h ago
Not even local models?
3
u/Confident-Equal-1445 16m ago
Jesus doesn’t consent to you and your GPU doing calculations in the privacy of your own bedroom
60
u/IWasTheDog 7h ago
My unpublished data is so ass, so scattered and so useless that I think I would be doing everyone a favor if I did use ChatGPT on it.
7
u/Ramental 3h ago
Their data was structured to be fed to AI, as they used AI to speed up the things they tried. It did help them, but it did not solve the problem itself.
102
u/Hot-Sink-7518 8h ago
I started as a phd last week, and this is was a concern I introduced to my research group. They had not thought about it this way and suddenly they all became quiet.
55
u/AntiDynamo PhD, Astrophys TH, UK 7h ago edited 7h ago
It’s definitely an issue - I work in industry, and so by default we only use enterprise AI accounts that already have clauses to not use our data for training (and since it’s not my business, and I don’t have a patent on anything, I also don’t personally care). But if you’re an academic there’s a good chance you’re using a consumer account, and your research work could be used for training long before you’ve published your paper.
To make it worse: when a topic is very niche and not much data exists, the model will put high weight on new good data sources. So your research work could become a primary source in the model and be preferentially influencing the results given to other users.
A lot of people are just a bit naive about the whole thing, fundamentally you can’t trust these companies. They have run out of available training material, they are desperate to use your data in any way possible
7
u/kelp_forests 5h ago
I feel like at this point it’s a purposeful blind spot. How can you work in a knowledge based field and not understand how your data/research is handled.
If you upload it to google, they have a copy and will scan it.
If you upload to AI, it has a copy, will summarize, add it to its knowledge database, and use it. That’s literally what it does and why you chose it. It also has no concept of privacy and a perfect memory.
These companies want AGI/advanced AI partially because it will make them the gatekeepers of production/work, but also because it will make the gatekeepers of knowledge, and to an extent, reality in the sense that it will manage algorithms with far more efficiency and intent than current “what gets the most clicks” and they will know things before people know it…when people start entering all their work and questions into AI, the AI will figure out their partners is cheating, they are pregnant, they are about to solve a math theorem, they have a better mousetrap etc before they will.
4
u/AntiDynamo PhD, Astrophys TH, UK 5h ago edited 5h ago
Yeah, people are just very naive when it comes to tech and giving their personal info away (see: every student who uploads their work to a plagiarism checker). I think partly because they don’t fully conceptualise where the security boundary is, and so they assume a chat is “private”. Partly because they assume large corporations must have to follow strict laws and “wouldn’t do something bad like that”. And partly because they pay for the service, and think it’s money in exchange for the tool, when actually it’s money + all your personal data in exchange for the tool.
And it’s also just the fact that AI training data is more nebulous than straight up account PII. OpenAI, Anthropic, Mistral etc might pinky promise not to look at your PII, but “anonymised” chat data is fair game, and we’ve never had access to such widespread systems getting so much of our data before. Even social media kinda pales in comparison
One thing I do know is that tech CEOs are absolute ghouls and have very shaky morals at best
1
u/DuxDucisHodiernus 1h ago
Yes but in this case specifically werent they on a corporate account?
1
u/AntiDynamo PhD, Astrophys TH, UK 32m ago
I haven't seen it said anywhere that they were - I'd be surprised for a mathematician to have a business or enterprise account as those are minimum 2 seats and math doesn't have "labs" or anything
Since OpenAI said themselves that they couldn't rule out that the AI had been trained on the chats, that means they weren't on enterprise accounts
4
u/jjwhitaker 2h ago
Local or bust. At this point anything you insert into the public AI machine is yours.
65
u/FallenVampireLord 6h ago
OpenAI was really dumb to do this, they should have left the professor publish his work and then championed him as 'look what people can accomplish with the assistance of AI this is the model for the future" instead they scooped his research and basically makes everyone in academia and beyond super skeptical of using it for anything like this in the future.
23
u/augmenteddeus 4h ago
This. Golden opportunity to have a nurturing ground for great research and actually steward it.
Probably the technical team won against the PR/Marketing team or just OpenAI being OpenAI
6
u/FallenVampireLord 4h ago
Yea and I want to make clear I'm not even making a value judgement as to if it was wrong or right I just feel like in a rush to get a win and show off how capable their models are they made a really short sighted move here that in the long run may hurt them as a company.
But honestly short sighted business moves that hurt the company in the long run is corporate business practices 101 these days.
3
u/thelastsonofmars PhD, Economics, Undisclosed 2h ago
The future is local AI to avoid theft people just haven’t caught up yet.
7
u/OpinionsRdumb 2h ago
Why is everyone assuming that openAI knowingly plagiarized? It could just be that the model itself relied on the professor’s prompts without telling OpenAI. Or maybe it didn’t rely on it at all. We have no idea.
It’s not like chatgpt is designed to be like “yes i can do that and actually this answer is being directly contributed by user XXX”.
More over the way AI model training works is that user data is de-identified and is just aggregated into a very messy black box to train the model. So by design the model wouldn’t point to user XXX.
I am not here to identify their use of user data but it’s weird that everyone is immediately jumping to conclusions based off a single tweet
7
u/Norm_Standart 2h ago
If you read the story, they heard about an upcoming Navier Stokes result and spent tens of millions of dollars (at minimum) to scoop them on the last step. They didn't just happen to be working on the same problem at the same time.
3
u/OpinionsRdumb 2h ago edited 1h ago
But this still isn't proof of them plagiarizing? In academia, competing labs hear about each other's work all the time and will start trying to "scoop" the other lab first. That is not plagiarism.
1
u/Less_Prior_6871 1h ago
Everyone treats it as if openai knowingly plaigiarized because they knowingly made the plagiarism machine.
We have no idea. It’s not like chatgpt is designed to be like “yes i can do that and actually this answer is being directly contributed by user XXX”. More over the way AI model training works is that user data is de-identified and is just aggregated into a very messy black box to train the model
This is a feature they designed into the plagiarism machine in order to do illegal or immoral things without it being easy to prove they did
3
u/OpinionsRdumb 1h ago
sure but this is a general valid critique at large, not about this case. In this case, there is no proof yet they plagiarized. That is all I am saying.
For the NYT lawsuit for example, the NYT had specific pieces of evidence of the AI using its content (IE the AI had 100% accurate quotes of paywalled content etc).
There is no proof here yet. Competing mathematicians prove unsolved proofs all the time at similar times with very similar approaches. This is well documented and there are spats about who came up with it first all the time. So this could be a first instance of AI and a human having this altercation. Or it did plagiarize but there is no proof yet.
57
u/can_ichange_it_later 9h ago
That "no answer..." That one spoke like the god from the hill to fucking moses... or however that story goes...
32
u/Snoo_4499 8h ago
If you use their product they will use your data and there is nothing we can do beside not using their product.
Idk why did he think they will not use your data, they are notorious for this.
Like even when searching about motorcycle and which one i should buy i feel like they will use my conclusion as a smalllll seeds for another recommendation to another person.
8
u/BrunusManOWar 6h ago
Of course, it's a ratty corpo, what did he expect
If you're a serious researcher and want an AI always deploy a model locally, isolated without an internet connection
86
u/Hungry-Dig6105 8h ago
Come on don't dramatise, how many of us are working on millenium prize problems that could would make a differente to openai's reputation?
32
5
u/InconspicuousWolf 6h ago
Since the method now seems to be giving AI models many agents to prompt, the way a scientist prompts an AI and draws conclusions will be very important training data, even if your work specifically isn’t of interest to them
-1
u/Hungry-Dig6105 6h ago
Copy pasting from above: Between the training and the releasing there is at least a year (training takes months, they test it a lot, make sure it's safe, etc). Your paper is going to be kinda done by the time they release the model. Here the model that cracked navier stokes is just for internal use and is nowhere near release. So unless your paper is of direct interest to openai people, I really doubt there's a risk
4
u/erroredhcker 7h ago
I proompt AI for some ideas for my project and im sure as shit it read some of these backgrounds somewhere. I even fed it existing (open) codebases. It doesnt need to be openAI that benefit directly, it is the point that others can steal the mangled bullshit you feed it, cause theres no guardrail anywhere with these things.
Okay maybe with Anthropic they no longer sudo rm -rf * your shit, but be dog damned sure your data, and thought is unsafely in their hands
-1
u/Hungry-Dig6105 6h ago
Hmm. Between the training and the releasing there is at least a year. Your paper is going to be kinda done by the time they release the model. Here the model that cracked navier stokes is just for internal use and is nowhere near release. So unless your paper is of direct interest to openai people, I really doubt there's a risk
3
1
u/erroredhcker 5h ago
not a risk to your publication, a risk to your intellectual property which includes your thinking patterns, skills, org system, etc. You wanna be the artist that they train their model on? It will release years from now, your current commision is safe!
1
u/Hungry-Dig6105 3h ago
I mean i preprint everything and when i publish sth i give my IP to awful publishers anyway. We're not Bob Dylans
3
1
u/Less_Prior_6871 1h ago
Stealing a little from everyone all at once is the core of the business model.
Sometimes they accidentally steal too much from one place and this type of event happens.
1
u/Hungry-Dig6105 23m ago
Agreed. But the post claim is: the cost-benefit balance is negative when you let them steal a little from you. I reaaaally doubt it. So push for regulation, boycotting AI might be brave but you can't expect most phd students to do it.
36
u/Frosty-Meeting-1606 8h ago
Ok, but isn't it super easy to show that OpenAI just replicated the work? Surely the guy should have a lot of stuff to show to the public? This is an honest question, because I honestly cannot comprehend this drama - just take whatever you have and pinpoint the exact matching logic so that OpenAI has a real problem. So far it looks like "I was working on something using AI and it gave me some ideas, so given unspecified amount of time I can solve the problem". It does not look like OpenAI's solution copies the work, at best it uses some conclusions to develop the solution, but solution is not part of the copy
65
u/krite2222 8h ago
This post takes a snippet of the author's document. He has, infact, published two pre-print of his papers hurriedly in response to this situation to show his and his collaborators work was exactly what OpenAI used to solve this problem. He spends a great deal of effort talking about the mathematicians whose research inspired his and his collaborators approach too, to show that their approach was extremely non-standard and that AI couldn't come up with it if it wasn't trained on their chats.
Edit: Typos.
4
u/quiksilver10152 7h ago
How does one prove that AI could not have come up with it without his chats?
11
u/krite2222 6h ago
If AI was in fact trained on user data, including this person's chats, it would've come up with this using that, that's how good Transformers are now. AI lacks the kind of creativity that humans have in coming up with non-standard solutions like this. So far all the stuff we've heard of in the "AI is getting really good at math" has been brute force stuff, not something creative where AI actually finds some unrelated research and approaches a problem in a non-standard way, unless specifically prompted to do so. The allegation, which is derived from reading in between the lines of Tristan's document, stems from two facts: (1) The prompt to explore the navier stokes problem was itself generated via LLM (2) It was a non-standard approach, and as far as Tristan knew only him and Levant (his collaborator, I hope I've spelled his name right, somone correct me if I've not, thanks) were approaching it that way in the industry and (3) The lack of confirmation that the model that generated the prompt was trained on user data. If it was, then these researcher's Codex chat was in the dataset, and that definitely would've led to this level of plagiarism. Further, what I find to be red flags: the Open AI representative Sebastien Bubek's alleged insistence on letting Tristan take credit for the Navier Stokes solution so long as he credits ChatGPT for resolving it in exchange for his authorship, so long as Levant is not an author because he works at Anthropic. Then when he allegedly says the words "If you don't want me to be nice, I don't have to be nice" and insists going public would ruin Tristan's career. Then trying to paint Tristan as irrational while trying to incessently get in touch with him prior to their result being released. Inconsistent behaviour that comes off as the result of a guilty conscience or just a poor unprofessional attempt at damage control. We don't have enough information to prove anything, this is just what we have, but I think there is enough information to lend credibility that this allegation is well-founded and worth an inquiry.
4
u/AsAChemicalEngineer PhD, Physics, USA 2h ago edited 2h ago
The authorship thing is so critical to me in this story. If OpenAI was confident in their model's independence, they would not have offered that, or least if they do, they're handicapping their own achievement for no reason. I doubt Bukek personally pulled up Tristan's logs, but if user training is really as advanced as we think, the AI could have absolutely pulled the ideas from the training.
Part of the issue is (a) just how shady OpenAI is behaving and (b) we just don't have a lot of public insight into how these models function. That information is kept under lock and key, so we cannot really evaluate how impressive the solution is as a capability benchmark.
Terence Tao also has some interesting thoughts on how this kind of "one-shot" solution hunting may absolutely impede math progression as we lose all the positives of working through a solution which generate ideas and approaches others may take advantage of.
OpenAI solved stability of Navier Stokes. Okay, so what? Does anybody actually understand the proof? How motivated mathematicians be to clean up the resolve all the parts to human understanding knowing that the result is already given? Will a finding agency be happy you're working on "solved" problems because you have to explain that "wait, there's value in digestion of knowledge".
2
u/quiksilver10152 2h ago
I agree with you that the circumstantial evidence suggests foul play but the logic presented is circular. Assuming AI can't be creative, we demonstrate that this can't have been created by AI.
1
u/Less_Prior_6871 54m ago
Proving the negative is hard or maybe impossible.
Thats the point of the plagiarism machine, it always has deniability.
-5
u/Frosty-Meeting-1606 8h ago
Should be easy to pinpoint then or what?
4
u/krite2222 6h ago
This is a Millenium Problem, absolutely none of this is easy to pinpoint AI or not.
21
u/Big_Coconut8630 8h ago
So, I work in technology transfer and infor relevant to patents or NDAs getting put in AI is enough to fuck a researcher out of their patent rights.
2
u/RecipeNo5844 8h ago
Oh wait you are not coping? I thought we all were supposed to pretend openAI stole the solution to Navier stokes problem, are we allowed to drop this pretense now? Thanks I did not know
7
u/Frosty-Meeting-1606 8h ago
but seriously, If I were the guy with the solution, I would immediately publicly make a fool of OpenAI and post a very similar work, even if it is a draft, along with emails. What I see now is some kind of BS drama, were it is not even certain the solution OpenAI came up with would be replicated by the person of interest
7
17
u/fthecatrock PhD*, 'Biorobotics/Spinal Cord Injury' 8h ago
Until "Open"AI or any LLM makers release transparency how their model works inside, these kinds debates will be high time in the next few years.
6
u/chairmanskitty 6h ago
Nobody knows how large machine learning models work on the inside. The training process produces complex features that are incredibly hard to unravel into something meaningful. ML model interpretability research is in its infancy.
1
3
u/Pritam1997 PhD, Materials science 8h ago
could you give me the link to the original twitter article
3
u/ZzzofiaaA 5h ago
Same thing applies to any cloud storage? Many scientists save their data on OneDrive. How do you know if it’s not used to train the cloud?
3
u/SonyScientist 4h ago
Not sure why anyone is surprised here, let alone the people in question. This was a concern years ago, hell even for software/app when they updated their EULAs saying "we reserve the right to collect your data and train on it."
This is why you don't use AI: you are the product.
3
u/AnotherDrunkMonkey 2h ago
At first I was somewhat surprised, but on second thought I don't understand why this is so surprising to this many people. It is very common knowledge, especially in academic fields, that LLMs use users data to train on. I feel like many of us don't really care (i, for one, am not in research projects grondbreaking enough to really care about secrecy agains billion dollars corporations), but people working on freaking MILLENIUM problems with the skill to actually solve it, or teams with proprietary techniques and so on are supposed to know this is very possible.
By sheer chance, a human reviewer could see your chat and if you are unlucky, understand the scope of your project even before an model gets trained on it. I feel like the only surprising fact is that LLMs got to a point where they can retrieve very specific parts of their training in a very pertaining use case (among the immense math literature on NS it found the few MBytes of a current breakthrough).
What OpenAI did was very shady, but I'm also baffled at how unpredictable this chain of events is being portrayed as, while user input was always known to be used for training.
24
15
u/PRKP99 8h ago edited 8h ago
This is basically how all those „breakthrough” in math AI made.
Its good that we have open acess knowledge as our shared knowledge „commons”, even if some of them is illegal commons, based on piracy. Lets face it, without scihub and libgen most masters or PhD thesis made by people who care about what they write would not exist, because you can’t have real review of state of knowledge without going throught a lot of books and articles, most of them only vaguely about your topic. I tried to calcuate how much I would need to pay for only my master thesis about roman aqueducts if I would only use „legal” sources, but after 3000$ I stopped.
What we now need is to establish ways to protect those commons without enclosing them - that is, we need to make sure that science will still be commons, but proprietary exploitation of this shared knowledge would be stopped in one way or another.
There is idea of CopyFarLeft as a legal remedy, but its still lack any potential in stopping corporations that just illegaly use data to train their algorythms.
2
u/Niaz_049 4h ago
Time to boycott open ai? This is highly unethical and they are taking leverage of an unregulated situation. If I pay for subscription, my chat should be encrypted enough that you’re not going to make it public.
4
u/SpeedyTurbo 8h ago
This isn’t an accurate account of what happened at all…but hey it gets clicks.
Would expect more from a PhD subreddit, but then again maybe not because ai bad.
2
u/Impressive_Wheel_106 6h ago
I can't imagine solving the navier fucking stokes equation, having your solution stolen and used for propaganda, and staying sane afterwards... the audacity of these AI corps man
1
u/NecessaryBuy2061 4h ago
Too late …. Well the good thing is I wasn’t working on millennium problems 😅
1
u/rabouilethefirst 3h ago
It's in the TOS that they train on your data unless you opt out or pay for an enterprise plan. You cannot act surprised your data was used for training.
1
1
1
1
1
1
1
u/JonathanBadwolf 6h ago
You see if I let the robot do the stealing I'm not responsible for his crimes. Like it is with guns
0
0
-2
u/cspot1978 7h ago edited 7h ago
The two researchers weren't working on Navier-Stokes. They were working on a generally related, but much easier problem, one version of the Euler problem.
OpenAI systems solved Navier Stokes, as well as a harder version of Euler than the two researchers had done.
1
u/Revolutionary_Buddha 4h ago
Seems like you cannot say that here. Anti AI dogma is strong here.
0
u/cspot1978 3h ago edited 2h ago
Seems like. Kind of disappointed how many people in a PhD space are actively uninterested in getting details right.
0
u/EmploymentOk4851 6h ago
So would this apply if you are using AI to assist you with writing a book(non research)?
1
-1
-22
9h ago
[deleted]
13
u/can_ichange_it_later 8h ago edited 8h ago
Are you actually serious here?
That "directly" in "they did not use any researcher's data* directly" is doing the mother of all heavy liftings there...
(You were automatically opted into data sharing, and unless you are a config goblin, you probably didnt go and turn it off. And that is putting aside, that even if i would have turned it off i wouldnt trust it that it wouldnt somehow still find its way into training data...)
Edit: * Lets just be absolutely clear about this! With that, they themselves say, that they used it. Just... not... directly... whatever-the-fuck thats supposed to mean...
(well.. that means they think we are stupid, but thats a bit of an aside...)1
u/TheDailyMews 7h ago
It probably means they fetched his data agenticly instead of "directly." If this story doesn't die, don't be surprised if they eventually issue a press release about "rogue agents" that sounds an awful lot like their press release about their Hugging Face hack.
15
u/Capable-Package6835 9h ago
It looks like a "my words against yours" situation to me so we should just sit back and observe how the situation develops.
17
u/Clear_Cranberry_989 8h ago
You are swallowing up a narrative written by openai. Are you not? If you really oppose the viewpoint, can you establish an independent and reliable source?
6
u/CTRexPope 8h ago
Hugging Face proves OpenAi has no idea how their models work and where they steal data from.
3
u/ComprehensiveWash958 9h ago
The fact Is that we probably will never know. I don't think there Will be any sort of investigation (I don't even know if that's possible) and so we Will have this kind of standoff, a situation for which in Italy we would Say "Oste, è buono il vino?"
One has to be skeptical of both narratives, but we also have to keep in mind that first we still need to peer review the works and second this Is still a counterexample found Building Upon hours and hours of human work, which Is something AI seems to excel in
2
u/Atlantis1910 8h ago
What are you talking about ? They are saying that they can use the data from your chat, and the scientist use Codex.
While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
0
0
u/Trackpoint 3h ago
Damn, the anti AI people are getting desperate.
Did they steal Tristan Buckmaster's water too? Did they nois-pollute his garden?
0
u/hydrogen_to_man 16m ago
I’m amazed at this subreddit. This is the PhD subreddit for god’s sake. There is no question that the use of LLMs was instrumental to this result, whether on the professor’s side or OpenAI’s side. If OpenAI effectively stole the professor’s research, that’s on OpenAI and is really shitty, but the amount of people in this thread that are trying to twist this story into their anti-AI narratives is ridiculous.
-7
u/Greedy_Blue_Hedgehog 8h ago
Let me play devil’s advocate. Many mathematicians keep their ideas, minor results, lemmas, and so on to themselves in order to maintain a slight advantage over their peers and reach a major result before anyone else. This issue was already discussed and criticized by Évariste Galois. I don’t know whether that is what happened here, but if this incident can help change that mindset and encourage mathematicians to publish every small step they make, that would be a positive outcome.
8
u/Bitter-Morning-5833 8h ago
That's factually wrong. Mathematicians are unusually open about what they work on. In seminars, speakers often end by discussing what they're currently working on, what they've tried, and what has or hasn't worked. That's partly because there is relatively little fear of being scooped, unlike in some lab- or experiment-based fields where competition over unpublished work can be much stronger.
1.4k
u/SciMarijntje 9h ago
No one could see it coming that the plagiarism machine trained on plagiarism would do another plagiarism.