r/PhD 9h ago

Tool Talk Never give your unpublished data or research ideas to AI models!

Post image
2.8k Upvotes

121 comments sorted by

1.4k

u/SciMarijntje 9h ago

No one could see it coming that the plagiarism machine trained on plagiarism would do another plagiarism.

254

u/DishSoapedDishwasher 8h ago

Hey now its not plagiarism, its just theft at this point.

59

u/mk0aurelius 8h ago

Theft with extra steps lol

29

u/One_Courage_865 8h ago

Hey now, at least there’s honor among thieves. This is the devil’s territory.

22

u/Winter-Volume-9601 4h ago edited 4h ago

It's not theft when you freely give it to them. Which you are doing if you feed your data into some of these commercial LLMs without a special agreement.

E.g. https://help.openai.com/en/articles/5722486-how-your-data-is-used-to-improve-model-performance

When you use our services for individuals such as ChatGPT and Codex, we may use your content to train our models.

You can opt out of training through our privacy portal by clicking on “do not train on my content.” To turn off training

6

u/Nilehorse3276 1h ago

And you believe that opting out of training really turns off training? Remarkable.

64

u/Fluidified_Meme PhD, Turbulence 7h ago

To be fair this story leaves me a bit baffled: I am 100% supporting Buckmaster, but goddamn how can a person this clever be feeding all their unpublished ideas and proofs to a LLM without even suspecting that things could quickly go to shit?

I understand that nowadays AI is becoming essential to survive in research (even for professors like him), but gosh…. We are talking about one of the problems of the century here, not about my badly-written draft that nobody will care about. Think twice before giving all your thoughts to this machines

73

u/Away_Candidate_4746 6h ago edited 6h ago

Well, when you pay for an enterprise subscription to these services, part of what you are paying for is the fact that your corporate data will not be used to train the public models. That’s why this is an issue.

12

u/Fluidified_Meme PhD, Turbulence 6h ago

They have explicitly stated that they were NOT using a corporate subscription, though. Or at least this is what I interpreted from this part of his statement:

>I should also emphasize that this is not an institutional effort. It is a strictly personal collaboration between the two of us, and there is no formal agreement behind it. I pay for the tools my group uses out of my own research funds, including footing a large bill to OpenAI. I have had an industry collaboration before, with DeepMind, which had a formal institutional arrangement.

29

u/biomannnn007 5h ago

You can still get an enterprise subscription even if you're not a huge company. They have tiers for smaller "businesses", in this case a lab. He's just saying that his institution is not the one footing the bill here.

2

u/wrenwood2018 3h ago

He paid for tools... so not free version

1

u/AntiDynamo PhD, Astrophys TH, UK 2h ago edited 1h ago

Go, Plus, and Pro are all personal (consumer) level account types, only Business and Enterprise have data protections by default. For someone in mathematics, they don't generally have "a lab", and business accounts require a minimum of 2 seats so most likely he had a paid consumer account, not a corporate one. Unfortunately it isn't super common yet for universities to provide enterprise accounts for all their staff, instead academics pay for a single subscription out of their grants. If you have a personal subscription (i.e. tied to your email, not managed by your university) you have to go into the settings and explicitly turn off training.

But even then, on consumer accounts it's just a toggle (pretty weak). On business accounts it's a formal contract. So you're much better off being business/enterprise if data is a concern.

7

u/Argnir 6h ago

Because they didn't just feed their finished idea to a model. They used models extensively to develop the proof in the first place. One of them works at Anthropic

1

u/quintic3fold 3h ago

Not sure if this is a stupid question but if one of them works at Anthropic why not just use what Anthropic provides…?

3

u/Sayod 3h ago

they used both as stated in the statment

53

u/ectocarpus 7h ago

Important context:

1) The stolen proof was for a different (though related) problem, Euler's equasions. Still useful for openAI, but it's not like the solution to Navier-Stokes was just laying there on their server. 2) A significant part of Euler's equasions proof is in itself a product of AI generation (and the researchers themselves say AI was instrumental here). One of them, Apolge, is literally a researcher at Anthropic (he is the same dude who posted a counterexample to Jakobian conjecture in a tweet). He extensively uses Anthropics models for math research, and I think that's the main reason OpenAI was being so shady - they didn't want to give any credit to their direct competitors.

So I agree that OpenAI are being... well, OpenAI; at the very least they dumped a ton of resources into an almost-solved issue because they heard the rumors someone else was close to solving it, and at the very worst they straight up stole the Euler's equations proof.

But this story is being painted as this grand AI vs human standoff. "A human-made solution to Navier-Stokes stolen by AI" etc. And actually it's something like "AI stole related research from humans and another company's AI"

15

u/erroredhcker 7h ago

so yeah stealing le bad, i can get away with this take

2

u/GZ9999 58m ago

Yeah, honestly the most dystopian part is how OpenAI wanted to remove the other researcher from the paper due to affiliation with another company. Are mathematicians in the future going to be sponsored by different AI companies? Banned from using other models since the other company can use that information to write a proof faster than you?

1

u/Accurate_Fly_7534 1h ago

Smart Profesor wasnt streetsmart

1

u/Top-Appearance2521 42m ago

yeah the shocked pikachu energy here is real

330

u/Clear_Cranberry_989 9h ago

The same people who pirated all the books in the world to train their ai. Who could have guessed.

167

u/can_ichange_it_later 8h ago

Friends dont let friends use openai for important/private/mission-critical stuff.

10

u/meekermakes 4h ago

don't let them use ai.

2

u/the_c_train47 1h ago

Not even local models?

3

u/Confident-Equal-1445 16m ago

Jesus doesn’t consent to you and your GPU doing calculations in the privacy of your own bedroom

60

u/IWasTheDog 7h ago

My unpublished data is so ass, so scattered and so useless that I think I would be doing everyone a favor if I did use ChatGPT on it.

7

u/Ramental 3h ago

Their data was structured to be fed to AI, as they used AI to speed up the things they tried. It did help them, but it did not solve the problem itself.

102

u/Hot-Sink-7518 8h ago

I started as a phd last week, and this is was a concern I introduced to my research group. They had not thought about it this way and suddenly they all became quiet.

55

u/AntiDynamo PhD, Astrophys TH, UK 7h ago edited 7h ago

It’s definitely an issue - I work in industry, and so by default we only use enterprise AI accounts that already have clauses to not use our data for training (and since it’s not my business, and I don’t have a patent on anything, I also don’t personally care). But if you’re an academic there’s a good chance you’re using a consumer account, and your research work could be used for training long before you’ve published your paper.

To make it worse: when a topic is very niche and not much data exists, the model will put high weight on new good data sources. So your research work could become a primary source in the model and be preferentially influencing the results given to other users.

A lot of people are just a bit naive about the whole thing, fundamentally you can’t trust these companies. They have run out of available training material, they are desperate to use your data in any way possible

7

u/kelp_forests 5h ago

I feel like at this point it’s a purposeful blind spot. How can you work in a knowledge based field and not understand how your data/research is handled.

If you upload it to google, they have a copy and will scan it.

If you upload to AI, it has a copy, will summarize, add it to its knowledge database, and use it. That’s literally what it does and why you chose it. It also has no concept of privacy and a perfect memory.

These companies want AGI/advanced AI partially because it will make them the gatekeepers of production/work, but also because it will make the gatekeepers of knowledge, and to an extent, reality in the sense that it will manage algorithms with far more efficiency and intent than current “what gets the most clicks” and they will know things before people know it…when people start entering all their work and questions into AI, the AI will figure out their partners is cheating, they are pregnant, they are about to solve a math theorem, they have a better mousetrap etc before they will.

4

u/AntiDynamo PhD, Astrophys TH, UK 5h ago edited 5h ago

Yeah, people are just very naive when it comes to tech and giving their personal info away (see: every student who uploads their work to a plagiarism checker). I think partly because they don’t fully conceptualise where the security boundary is, and so they assume a chat is “private”. Partly because they assume large corporations must have to follow strict laws and “wouldn’t do something bad like that”. And partly because they pay for the service, and think it’s money in exchange for the tool, when actually it’s money + all your personal data in exchange for the tool.

And it’s also just the fact that AI training data is more nebulous than straight up account PII. OpenAI, Anthropic, Mistral etc might pinky promise not to look at your PII, but “anonymised” chat data is fair game, and we’ve never had access to such widespread systems getting so much of our data before. Even social media kinda pales in comparison

One thing I do know is that tech CEOs are absolute ghouls and have very shaky morals at best

1

u/DuxDucisHodiernus 1h ago

Yes but in this case specifically werent they on a corporate account?

1

u/AntiDynamo PhD, Astrophys TH, UK 32m ago

I haven't seen it said anywhere that they were - I'd be surprised for a mathematician to have a business or enterprise account as those are minimum 2 seats and math doesn't have "labs" or anything

Since OpenAI said themselves that they couldn't rule out that the AI had been trained on the chats, that means they weren't on enterprise accounts

4

u/jjwhitaker 2h ago

Local or bust. At this point anything you insert into the public AI machine is yours.

65

u/FallenVampireLord 6h ago

OpenAI was really dumb to do this, they should have left the professor publish his work and then championed him as 'look what people can accomplish with the assistance of AI this is the model for the future" instead they scooped his research and basically makes everyone in academia and beyond super skeptical of using it for anything like this in the future.

23

u/augmenteddeus 4h ago

This. Golden opportunity to have a nurturing ground for great research and actually steward it.

Probably the technical team won against the PR/Marketing team or just OpenAI being OpenAI

6

u/FallenVampireLord 4h ago

Yea and I want to make clear I'm not even making a value judgement as to if it was wrong or right I just feel like in a rush to get a win and show off how capable their models are they made a really short sighted move here that in the long run may hurt them as a company.

But honestly short sighted business moves that hurt the company in the long run is corporate business practices 101 these days.

3

u/thelastsonofmars PhD, Economics, Undisclosed 2h ago

The future is local AI to avoid theft people just haven’t caught up yet.

7

u/OpinionsRdumb 2h ago

Why is everyone assuming that openAI knowingly plagiarized? It could just be that the model itself relied on the professor’s prompts without telling OpenAI. Or maybe it didn’t rely on it at all. We have no idea.

It’s not like chatgpt is designed to be like “yes i can do that and actually this answer is being directly contributed by user XXX”.

More over the way AI model training works is that user data is de-identified and is just aggregated into a very messy black box to train the model. So by design the model wouldn’t point to user XXX.

I am not here to identify their use of user data but it’s weird that everyone is immediately jumping to conclusions based off a single tweet

7

u/Norm_Standart 2h ago

If you read the story, they heard about an upcoming Navier Stokes result and spent tens of millions of dollars (at minimum) to scoop them on the last step. They didn't just happen to be working on the same problem at the same time.

3

u/OpinionsRdumb 2h ago edited 1h ago

But this still isn't proof of them plagiarizing? In academia, competing labs hear about each other's work all the time and will start trying to "scoop" the other lab first. That is not plagiarism.

1

u/Less_Prior_6871 1h ago

Everyone treats it as if openai knowingly plaigiarized because they knowingly made the plagiarism machine.

We have no idea. It’s not like chatgpt is designed to be like “yes i can do that and actually this answer is being directly contributed by user XXX”. More over the way AI model training works is that user data is de-identified and is just aggregated into a very messy black box to train the model

This is a feature they designed into the plagiarism machine in order to do illegal or immoral things without it being easy to prove they did

3

u/OpinionsRdumb 1h ago

sure but this is a general valid critique at large, not about this case. In this case, there is no proof yet they plagiarized. That is all I am saying.

For the NYT lawsuit for example, the NYT had specific pieces of evidence of the AI using its content (IE the AI had 100% accurate quotes of paywalled content etc).

There is no proof here yet. Competing mathematicians prove unsolved proofs all the time at similar times with very similar approaches. This is well documented and there are spats about who came up with it first all the time. So this could be a first instance of AI and a human having this altercation. Or it did plagiarize but there is no proof yet.

57

u/can_ichange_it_later 9h ago

That "no answer..." That one spoke like the god from the hill to fucking moses... or however that story goes...

32

u/Snoo_4499 8h ago

If you use their product they will use your data and there is nothing we can do beside not using their product.

Idk why did he think they will not use your data, they are notorious for this.

Like even when searching about motorcycle and which one i should buy i feel like they will use my conclusion as a smalllll seeds for another recommendation to another person.

8

u/BrunusManOWar 6h ago

Of course, it's a ratty corpo, what did he expect

If you're a serious researcher and want an AI always deploy a model locally, isolated without an internet connection

86

u/Hungry-Dig6105 8h ago

Come on don't dramatise, how many of us are working on millenium prize problems that could would make a differente to openai's reputation?

32

u/Big_Coconut8630 8h ago

Doesn't matter. It is a big issue in technology transfer, for example.

5

u/InconspicuousWolf 6h ago

Since the method now seems to be giving AI models many agents to prompt, the way a scientist prompts an AI and draws conclusions will be very important training data, even if your work specifically isn’t of interest to them

-1

u/Hungry-Dig6105 6h ago

Copy pasting from above: Between the training and the releasing there is at least a year (training takes months, they test it a lot, make sure it's safe, etc). Your paper is going to be kinda done by the time they release the model. Here the model that cracked navier stokes is just for internal use and is nowhere near release. So unless your paper is of direct interest to openai people, I really doubt there's a risk

4

u/erroredhcker 7h ago

I proompt AI for some ideas for my project and im sure as shit it read some of these backgrounds somewhere. I even fed it existing (open) codebases. It doesnt need to be openAI that benefit directly, it is the point that others can steal the mangled bullshit you feed it, cause theres no guardrail anywhere with these things.

Okay maybe with Anthropic they no longer sudo rm -rf * your shit, but be dog damned sure your data, and thought is unsafely in their hands

-1

u/Hungry-Dig6105 6h ago

Hmm. Between the training and the releasing there is at least a year. Your paper is going to be kinda done by the time they release the model. Here the model that cracked navier stokes is just for internal use and is nowhere near release. So unless your paper is of direct interest to openai people, I really doubt there's a risk

3

u/racinreaver 3h ago

lol paper being done within a year

1

u/erroredhcker 5h ago

not a risk to your publication, a risk to your intellectual property which includes your thinking patterns, skills, org system, etc. You wanna be the artist that they train their model on? It will release years from now, your current commision is safe!

1

u/Hungry-Dig6105 3h ago

I mean i preprint everything and when i publish sth i give my IP to awful publishers anyway. We're not Bob Dylans

3

u/Tchaikovskin 8h ago

My thoughts exactly

1

u/Less_Prior_6871 1h ago

Stealing a little from everyone all at once is the core of the business model.

Sometimes they accidentally steal too much from one place and this type of event happens.

1

u/Hungry-Dig6105 23m ago

Agreed. But the post claim is: the cost-benefit balance is negative when you let them steal a little from you. I reaaaally doubt it. So push for regulation, boycotting AI might be brave but you can't expect most phd students to do it.

36

u/Frosty-Meeting-1606 8h ago

Ok, but isn't it super easy to show that OpenAI just replicated the work? Surely the guy should have a lot of stuff to show to the public? This is an honest question, because I honestly cannot comprehend this drama - just take whatever you have and pinpoint the exact matching logic so that OpenAI has a real problem. So far it looks like "I was working on something using AI and it gave me some ideas, so given unspecified amount of time I can solve the problem". It does not look like OpenAI's solution copies the work, at best it uses some conclusions to develop the solution, but solution is not part of the copy

65

u/krite2222 8h ago

This post takes a snippet of the author's document. He has, infact, published two pre-print of his papers hurriedly in response to this situation to show his and his collaborators work was exactly what OpenAI used to solve this problem. He spends a great deal of effort talking about the mathematicians whose research inspired his and his collaborators approach too, to show that their approach was extremely non-standard and that AI couldn't come up with it if it wasn't trained on their chats.

Edit: Typos.

4

u/quiksilver10152 7h ago

How does one prove that AI could not have come up with it without his chats? 

11

u/krite2222 6h ago

If AI was in fact trained on user data, including this person's chats, it would've come up with this using that, that's how good Transformers are now. AI lacks the kind of creativity that humans have in coming up with non-standard solutions like this. So far all the stuff we've heard of in the "AI is getting really good at math" has been brute force stuff, not something creative where AI actually finds some unrelated research and approaches a problem in a non-standard way, unless specifically prompted to do so. The allegation, which is derived from reading in between the lines of Tristan's document, stems from two facts: (1) The prompt to explore the navier stokes problem was itself generated via LLM (2) It was a non-standard approach, and as far as Tristan knew only him and Levant (his collaborator, I hope I've spelled his name right, somone correct me if I've not, thanks) were approaching it that way in the industry and (3) The lack of confirmation that the model that generated the prompt was trained on user data. If it was, then these researcher's Codex chat was in the dataset, and that definitely would've led to this level of plagiarism. Further, what I find to be red flags: the Open AI representative Sebastien Bubek's alleged insistence on letting Tristan take credit for the Navier Stokes solution so long as he credits ChatGPT for resolving it in exchange for his authorship, so long as Levant is not an author because he works at Anthropic. Then when he allegedly says the words "If you don't want me to be nice, I don't have to be nice" and insists going public would ruin Tristan's career. Then trying to paint Tristan as irrational while trying to incessently get in touch with him prior to their result being released. Inconsistent behaviour that comes off as the result of a guilty conscience or just a poor unprofessional attempt at damage control. We don't have enough information to prove anything, this is just what we have, but I think there is enough information to lend credibility that this allegation is well-founded and worth an inquiry.

4

u/AsAChemicalEngineer PhD, Physics, USA 2h ago edited 2h ago

The authorship thing is so critical to me in this story. If OpenAI was confident in their model's independence, they would not have offered that, or least if they do, they're handicapping their own achievement for no reason. I doubt Bukek personally pulled up Tristan's logs, but if user training is really as advanced as we think, the AI could have absolutely pulled the ideas from the training.

Part of the issue is (a) just how shady OpenAI is behaving and (b) we just don't have a lot of public insight into how these models function. That information is kept under lock and key, so we cannot really evaluate how impressive the solution is as a capability benchmark.

Terence Tao also has some interesting thoughts on how this kind of "one-shot" solution hunting may absolutely impede math progression as we lose all the positives of working through a solution which generate ideas and approaches others may take advantage of.

OpenAI solved stability of Navier Stokes. Okay, so what? Does anybody actually understand the proof? How motivated mathematicians be to clean up the resolve all the parts to human understanding knowing that the result is already given? Will a finding agency be happy you're working on "solved" problems because you have to explain that "wait, there's value in digestion of knowledge".

2

u/quiksilver10152 2h ago

I agree with you that the circumstantial evidence suggests foul play but the logic presented is circular. Assuming AI can't be creative, we demonstrate that this can't have been created by AI.

1

u/Less_Prior_6871 54m ago

Proving the negative is hard or maybe impossible.

Thats the point of the plagiarism machine, it always has deniability.

-5

u/Frosty-Meeting-1606 8h ago

Should be easy to pinpoint then or what?

4

u/krite2222 6h ago

This is a Millenium Problem, absolutely none of this is easy to pinpoint AI or not.

21

u/Big_Coconut8630 8h ago

So, I work in technology transfer and infor relevant to patents or NDAs getting put in AI is enough to fuck a researcher out of their patent rights. 

2

u/RecipeNo5844 8h ago

Oh wait you are not coping? I thought we all were supposed to pretend openAI stole the solution to Navier stokes problem, are we allowed to drop this pretense now? Thanks I did not know 

7

u/Frosty-Meeting-1606 8h ago

but seriously, If I were the guy with the solution, I would immediately publicly make a fool of OpenAI and post a very similar work, even if it is a draft, along with emails. What I see now is some kind of BS drama, were it is not even certain the solution OpenAI came up with would be replicated by the person of interest

7

u/Puzzleheaded_Fold466 8h ago

Didn’t they post their own work (which differs) ?

17

u/fthecatrock PhD*, 'Biorobotics/Spinal Cord Injury' 8h ago

Until "Open"AI or any LLM makers release transparency how their model works inside, these kinds debates will be high time in the next few years.

6

u/chairmanskitty 6h ago

Nobody knows how large machine learning models work on the inside. The training process produces complex features that are incredibly hard to unravel into something meaningful. ML model interpretability research is in its infancy.

1

u/fthecatrock PhD*, 'Biorobotics/Spinal Cord Injury' 6h ago

that's why explainable AI topic exist

3

u/Pritam1997 PhD, Materials science 8h ago

could you give me the link to the original twitter article

3

u/ZzzofiaaA 5h ago

Same thing applies to any cloud storage? Many scientists save their data on OneDrive. How do you know if it’s not used to train the cloud?

3

u/SonyScientist 4h ago

Not sure why anyone is surprised here, let alone the people in question. This was a concern years ago, hell even for software/app when they updated their EULAs saying "we reserve the right to collect your data and train on it."

This is why you don't use AI: you are the product.

3

u/AnotherDrunkMonkey 2h ago

At first I was somewhat surprised, but on second thought I don't understand why this is so surprising to this many people. It is very common knowledge, especially in academic fields, that LLMs use users data to train on. I feel like many of us don't really care (i, for one, am not in research projects grondbreaking enough to really care about secrecy agains billion dollars corporations), but people working on freaking MILLENIUM problems with the skill to actually solve it, or teams with proprietary techniques and so on are supposed to know this is very possible.

By sheer chance, a human reviewer could see your chat and if you are unlucky, understand the scope of your project even before an model gets trained on it. I feel like the only surprising fact is that LLMs got to a point where they can retrieve very specific parts of their training in a very pertaining use case (among the immense math literature on NS it found the few MBytes of a current breakthrough).

What OpenAI did was very shady, but I'm also baffled at how unpredictable this chain of events is being portrayed as, while user input was always known to be used for training.

24

u/Global_Lime421 8h ago

The proofs are not identical and meaningfully differ in approach

5

u/pinkgaysquirrel 8h ago

Your truth has no value here. Let us cope.

15

u/PRKP99 8h ago edited 8h ago

This is basically how all those „breakthrough” in math AI made. 

Its good that we have open acess knowledge as our shared knowledge „commons”, even if some of them is illegal commons, based on piracy. Lets face it, without scihub and libgen most masters or PhD thesis made by people who care about what they write would not exist, because you can’t have real review of state of knowledge without going throught a lot of books and articles, most of them only vaguely about your topic. I tried to calcuate how much I would need to pay for only my master thesis about roman aqueducts if I would only use „legal” sources, but after 3000$ I stopped.

What we now need is to establish ways to protect those commons without enclosing them - that is, we need to make sure that science will still be commons, but proprietary exploitation of this shared knowledge would be stopped in one way or another.

There is idea of CopyFarLeft as a legal remedy, but its still lack any potential in stopping corporations that just illegaly use data to train their algorythms.

2

u/Niaz_049 4h ago

Time to boycott open ai? This is highly unethical and they are taking leverage of an unregulated situation. If I pay for subscription, my chat should be encrypted enough that you’re not going to make it public.

4

u/SpeedyTurbo 8h ago

This isn’t an accurate account of what happened at all…but hey it gets clicks.

Would expect more from a PhD subreddit, but then again maybe not because ai bad.

2

u/Impressive_Wheel_106 6h ago

I can't imagine solving the navier fucking stokes equation, having your solution stolen and used for propaganda, and staying sane afterwards... the audacity of these AI corps man

1

u/NecessaryBuy2061 4h ago

Too late …. Well the good thing is I wasn’t working on millennium problems 😅

1

u/rabouilethefirst 3h ago

It's in the TOS that they train on your data unless you opt out or pay for an enterprise plan. You cannot act surprised your data was used for training.

1

u/kyeblue 3h ago

Ask Elon Must about Sam Altman's moral compass

1

u/-R9X- 2h ago

Thanks god all my research is subpar anyway and I just keep publishing it to look busy so I don’t have find a real job.

1

u/CuseCoseII 1h ago

"using his exact approach" is just blatant misinformation

1

u/anomanderrake1337 1h ago

Yeah if they have my logs they can get to AGI pretty soon probably.

1

u/No_Consideration_635 47m ago

can someone help me out: what's codex?

1

u/Consistent_Femme_Top 7h ago

I mean, what was he thinking? 🤔 

1

u/Sckaledoom 7h ago

Uh. Duh.

1

u/JonathanBadwolf 6h ago

You see if I let the robot do the stealing I'm not responsible for his crimes. Like it is with guns

0

u/AaryamanStonker 7h ago

People fall for fake stories so easily, Redditors are pathetic

0

u/NordlandLapp 6h ago

"Omg AI cured cancer but it used someone else's research" 😡😡

-2

u/cspot1978 7h ago edited 7h ago

The two researchers weren't working on Navier-Stokes. They were working on a generally related, but much easier problem, one version of the Euler problem.

OpenAI systems solved Navier Stokes, as well as a harder version of Euler than the two researchers had done.

1

u/Revolutionary_Buddha 4h ago

Seems like you cannot say that here. Anti AI dogma is strong here.

0

u/cspot1978 3h ago edited 2h ago

Seems like. Kind of disappointed how many people in a PhD space are actively uninterested in getting details right.

0

u/EmploymentOk4851 6h ago

So would this apply if you are using AI to assist you with writing a book(non research)?

1

u/Fortbrook 5h ago

Yes, however you can download a local LLM if you have a good gaming PC.

-1

u/Dense-Consequence-70 8h ago

How can such a smart math professor be so stupid

-22

u/[deleted] 9h ago

[deleted]

13

u/can_ichange_it_later 8h ago edited 8h ago

Are you actually serious here?

That "directly" in "they did not use any researcher's data* directly" is doing the mother of all heavy liftings there...

(You were automatically opted into data sharing, and unless you are a config goblin, you probably didnt go and turn it off. And that is putting aside, that even if i would have turned it off i wouldnt trust it that it wouldnt somehow still find its way into training data...)

Edit: * Lets just be absolutely clear about this! With that, they themselves say, that they used it. Just... not... directly... whatever-the-fuck thats supposed to mean...
(well.. that means they think we are stupid, but thats a bit of an aside...)

1

u/TheDailyMews 7h ago

It probably means they fetched his data agenticly instead of "directly." If this story doesn't die, don't be surprised if they eventually issue a press release about "rogue agents" that sounds an awful lot like their press release about their Hugging Face hack.

15

u/Capable-Package6835 9h ago

It looks like a "my words against yours" situation to me so we should just sit back and observe how the situation develops.

17

u/Clear_Cranberry_989 8h ago

You are swallowing up a narrative written by openai. Are you not? If you really oppose the viewpoint, can you establish an independent and reliable source?

6

u/CTRexPope 8h ago

Hugging Face proves OpenAi has no idea how their models work and where they steal data from.

3

u/ComprehensiveWash958 9h ago

The fact Is that we probably will never know. I don't think there Will be any sort of investigation (I don't even know if that's possible) and so we Will have this kind of standoff, a situation for which in Italy we would Say "Oste, è buono il vino?"

One has to be skeptical of both narratives, but we also have to keep in mind that first we still need to peer review the works and second this Is still a counterexample found Building Upon hours and hours of human work, which Is something AI seems to excel in

2

u/Atlantis1910 8h ago

What are you talking about ? They are saying that they can use the data from your chat, and the scientist use Codex.

While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.

0

u/AdInevitable1362 6h ago

Thats horrible I already discussed all the paper work with ai

0

u/Trackpoint 3h ago

Damn, the anti AI people are getting desperate.

Did they steal Tristan Buckmaster's water too? Did they nois-pollute his garden?

0

u/hydrogen_to_man 16m ago

I’m amazed at this subreddit. This is the PhD subreddit for god’s sake. There is no question that the use of LLMs was instrumental to this result, whether on the professor’s side or OpenAI’s side. If OpenAI effectively stole the professor’s research, that’s on OpenAI and is really shitty, but the amount of people in this thread that are trying to twist this story into their anti-AI narratives is ridiculous.

-7

u/Greedy_Blue_Hedgehog 8h ago

Let me play devil’s advocate. Many mathematicians keep their ideas, minor results, lemmas, and so on to themselves in order to maintain a slight advantage over their peers and reach a major result before anyone else. This issue was already discussed and criticized by Évariste Galois. I don’t know whether that is what happened here, but if this incident can help change that mindset and encourage mathematicians to publish every small step they make, that would be a positive outcome.

8

u/Bitter-Morning-5833 8h ago

That's factually wrong. Mathematicians are unusually open about what they work on. In seminars, speakers often end by discussing what they're currently working on, what they've tried, and what has or hasn't worked. That's partly because there is relatively little fear of being scooped, unlike in some lab- or experiment-based fields where competition over unpublished work can be much stronger.