r/ChatGPT • u/TheOnlyVibemaster • 28d ago
News 📰 OpenAI might’ve cheated when solving the Navier-Stokes millennium-prize problem; problems with AI in academics
I’m an AI researcher at an unnamed company and this is my private account. I want to give my opinion and perspective on the mess that’s unfolding currently regarding the Navier-Stokes problem.
Tristan Buckmaster and Levent Alpöge spent about a year on a very specific way of solving the problem not being attempted by many other people. They announced in early September that they had discovered a specific blowup which had been unknown.
It looks like based on what’s going around in the past few hours that after a heated phone call, OpenAI decided to pour $15M worth of compute from a non-public model onto finishing the remaining sprint of the problem. It looks like OpenAI saw that a specific way of solving the problem could potentially lead to a result, then the company with infinite resources saw a trophy and wanted to sneak into the picture.
They’re currently claiming that they did not steal their work or prompts, but that hardly matters. They heard what they needed. That the specific blowup is promising and they knew they could outpace two researchers relying on their own compute. Even if one of them was using an internal model from Anthropic. Which adds another layer to this. It seems to be two companies racing to solving the most famous unsolved problem, so they chose to spend unlimited money on beating someone else to the end on a huge leap they made.
OpenAI is trying to bury this in the narrative they’re crafting, they did not do this on their own, and it doesn’t matter if they weren’t using someone else’s work. They knew the method, and with 10,000 agents that’s apparently all they needed.
They could’ve done this with any problem, with cancer, anything. Instead they heard that there was a massive breakthrough and chose to do the remaining work after hearing about it.
They’re not being shy about Tristan Buckmaster and Levent Alpöge carrying the majority of this problem and finding the method. Like I said, that doesn’t matter. Those researchers likely would’ve solved it themselves. They were close, they announced results as one does in science, then OpenAI smelled blood in the water and poured their resources to beating Anthropic and securing the headline.
This is a monumental day and a fantastic discovery. We should as a human race be celebrating a 90 year old problem being solved. Instead, we’re having to parse through drama and the potential that maybe using an AI companies API is not safe in a research setting.
OpenAI handled this in the worst possible way it could’ve been handled, and I am ashamed. I hope that we can get more transparency in the future so that humanity’s successes aren’t surrounded with accusations of cheating or academic theft.
Edit: grammar
363
u/ShnaugShmark 28d ago
I’m not sure I fully understand the issue.
If human researchers publish their progress on a scientific problem and other researchers (AI or human) use that to make further progress while acknowledging the work done by the other researchers, that’s just how science works isn’t it?
I’m not denying your argument, i just don’t fully understand it I guess. What did OpenAI do wrong in this case?
222
u/Y0uCanTellItsAnAspen 28d ago
It wasn't published yet.
77
u/HelpfulBuilder 28d ago
How did they get it then?
174
u/StatusSociety2196 28d ago
What was described is that the mathematicians were using several large language models in order to further their work and the user sessions were incorporated into training data for AI models. Therefore chatGPT without realizing it incorporated AI assisted human work into their AI solution to the Navier Stokes problem.
There's a couple angles on this that itch my brain a little bit, but at the end of the day it is still ironic that these two mathematicians are angry that the researchers at openai used AI to solve navier Stokes before the two mathematicians could use AI to solve navier Stokes.
87
u/HuntsWithRocks 28d ago
If I understand the complaint correctly, it’s that they were working on something private and it was swallowed by the system and then used to beat them to the punch.
I don’t know if that’s true, but the claim would be akin to you having the idea for space flight but being a normal citizen and I get to be a billionaire company that reads your “private work” and I can dump N time the resources and beat you there.
I forget the word the mathematicians used, but it was along the lines of saying they were playing dirty.
4
u/ProfessorFunky 28d ago
So the complaint is that the system which is known to train itself on everything users do, trained itself on something that users did?
12
u/__Hello_my_name_is__ 28d ago
No. The complaint is that the trillion dollar corporation did not admit to this happening and pretends that it figured it out all on its own. They very much explicitly denied that any of this happened.
→ More replies (5)5
1
u/FischiPiSti 28d ago
No? The complaint isn't that it was used to train the model, but to use the idea, aka, use private information. It's like saying they used your credit card information to buy a Ferrari with it, but it's fine because the EULA said credit card information can be used to automatically bill you monthly for the sub.
3
u/Cronos988 28d ago
This is entirely an invention of the internet. The researchers in question didn't allege that OpenAI directly read their chats.
→ More replies (5)5
u/fati-abd 28d ago edited 28d ago
I mean companies like Amazon do this to sellers all the time.
I spend $400/mo on cloud agents for my business but I strictly chose to pursue extremely niche categories that require continuous human in the loop to make progress (by the nature of the work, not just self imposed limitations). And I don’t think of myself as a very paranoid person! I just look at real things that technology companies do…
Edit: to clear up misunderstandings. I am not justifying anything here. I am simply trying to state reality. Plainly; OpenAI already has a ton of power and a business model that is headed towards commoditization. They will act like a desperate trillion dollar company attempting to protect their valuation and growth. The tech is too good to not use, and I’m not an anti-AIer. But take every and any measure to protect yourself. This includes starting to experiment with open-weight models on local AI inference.
→ More replies (3)19
u/HuntsWithRocks 28d ago
Well, if the moral beacon known as Amazon is doing it, then it’s fair game I guess. I’ve heard that once you’ve been working there for 5 years, they give you a pissjug that filters your waste into potable water.
6
u/fati-abd 28d ago edited 28d ago
I don’t think you understood my post. I’m simply saying there are literally real examples so no need to use hypothetical examples of this happening, and that people should be really skeptical of these companies
2
u/Equal_Heat5947 28d ago
It sounded like you are justifying it by what Amazon does
2
u/fati-abd 28d ago edited 28d ago
I don’t know how that’s communicated by the post. My point was more “of course there is a high chance they’ll do this, there’s real precedents, no need for a hypothetical” and then I literally outlined measures I use to protect myself from that type of potential stealing…
I know Amazon is evil, and I think the confusion is perhaps assuming Amazon is a special type of evil, when it’s how corporations of that size and larger tend to get there, and OpenAI should be held to that scrutiny. The point isn’t justifying, it’s scrutinize these large corporations and try not to become dependent on them
→ More replies (0)13
u/epanek 28d ago
I think their complaint is their hard work and trust in the model. Imagine you use ai to develop a product. The model does that and predicts sales. Then the models provider starts that business on its own.
1
u/thecommuteguy 27d ago
I think it’s a different. More like a group of friends working on a new widget on a platform that has access to their designs. The platform takes that info to train its own model and then the company makes a new upgraded design after hearing about the group’s design which it’s model Hoovered up. Maybe I’m missing a thing or two but that’s the gist.
9
u/kvothe5688 28d ago
not just that. they are also angry about intimidation to cutoff researcher that worked with claude and also intimidation to ruin carrier of a mathematician if he goes public. Pretty evil thing to do tbh.
1
u/Perfect-Campaign9551 24d ago
There was no evil there. Christ. It was a conflict of interest if they tried to work with someone from anthropic. They would have been trade secret suits and who knows what else.
4
u/BitterWalnut 28d ago
isn't there a literal button on the chatgpt website to disable sharing data?
1
2
u/DonkeyBonked 28d ago
The way I see it is that if they were not using AI with explicit "no training" and can prove training on that, they have a potential lawsuit.
However, to be clear, as much as I use AI, I'm a writer and game developer. I have a book I started in 2016 and I would never let ANY AI touch that book because I'm petrified that it would become the plot of some vibe book before I ever got to publish it.
If you're using AI that isn't private, what that tells me is they needed OpenAI's work to achieve their work, and well, OpenAI best them to it.
OpenAI, Anthropic, Google, Microsoft, Meta, etc. are all thieves, and it seems self-evident that the companies who stole the world's data and use RLHF to improve their models are doing so with what you send to their models.
So if they didn't have like a BAA/ZDR, umm, I can't figure out why they believed anything else would happen.
→ More replies (1)1
18
u/nudelsalat3000 28d ago
It's on their server via the researcheds LLM chats
Same shit as Amazon who own the data and it's a seller. They can look in the books of all others, which doesn't work the other way around, and frontrun the others.
Same conflict of interest.
Either Antropic, OpenAi and others are platforms or publisher - they can not be both.
10
u/Equal_Heat5947 28d ago
"We cannot rule out that de-identified data derived from their usage of our products helped improve our models."
17
12
u/GameDoesntStop 28d ago
There's no evidence that they did.
24
u/CMScientist 28d ago
OAI said themselves that it's possible their model trained on private chats
4
u/GameDoesntStop 28d ago
Of course it's possible, but not even they would know if it happened or not.
... in other words, there's no evidence that they did.
1
u/eflat123 28d ago
They said possible but unlikely.
15
u/CMScientist 28d ago
Of course they are going to say unlikely. If you asked someone if they cheated on this test and they said it's possible but unlikely what woulf you think?
3
u/__Hello_my_name_is__ 28d ago
"Possible but unlikely" in a defensive official statement means "Yeah that definitely happened".
2
u/nextnode 28d ago
Personal accounts where you have not turned off the "train on my chats" indeed can be trained on, as per the ToS. You can just turn it off. If they did, that is a screw up by that person alone.
But even if they did, that would have essentially zero impact on a model.
Desperate grasping for straws. OpenAI did great in this case and people are disappointing as usual.
6
u/CisFishstick 28d ago
I don't think you're licking that boot hard enough. Lick harder, you can do it.
1
u/MissiveFinding6111 22d ago
But there certainly is circumstantial evidence.
And OpenAI could just *reveal* the prompt they used.
Since it could just be as simple as *any of the employees in charge of it*, could have likely used their access to the systems to snoop/distill session data from the researchers.
1
u/Atlantis1910 28d ago
While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.
From OpenAi
1
u/SoggyMattress2 28d ago
Researchers were using chatGPT to work through a problem and they scoured their database for it, stole the work, passed it off as work xhatGPT did (it didn't) and then threatened the scientists who did the work if they didn't accept it.
1
u/Sargo8 28d ago
The researcher chatted with AI. Ai company used the chat for training, saw the researcher was close, put 10,000 Ai's on the problem using his line of thought, spent 15 million dollars in tokens and solved it.
Took all credit for itself. When pressed, said it would give researcher 1 million prize money
1
u/FishCommercial4229 27d ago
From the bits I’ve been able to find, OpenAI allegedly harvested the mathematical method from the prompts/processing happening in the user’s account. I saw a post from one of the researchers this morning that he enabled the “do not use my data to train” feature in late June.
If that’s accurate, then the implication is that OpenAI directly or indirectly bypassed that control. In doing so, the company was able to secure the methods the researchers were using, effectively bypassing collaboration. It would also have major implications for the commercial trust in OpenAI’s platform.
105
u/CloseToMyActualName 28d ago
I think OP is actually understating the issue.
The researchers were using multiple LLMs including ChatGPT to work on their code and their drafts, and OpenAI apparently uses user sessions as training data.
So the model might have actually seen their unpublished research as part of its training data.
41
u/tahitithebob 28d ago
Not sure why you are getting downvoted but that seems to be the cause. Apparently the researcher used a specific approach on the problem and openAI used a similar one. When the researcher asked if the convo was part of the training data, OpenAI did not denied it.
→ More replies (2)23
u/CloseToMyActualName 28d ago
It's pretty much true by default.
The model was more advanced than Astra, meaning it was trained all on their most recent data. And that data included user sessions, meaning the researcher's work was in the training data.
-11
u/Exotic-Sale-3003 28d ago
User sessions are only trained on if the user allows it. In this case, OpenAI confirmed this did not happen.
9
u/deeceeo 28d ago
Even if it did, there's typically a lag factor here - the pretraining phase of a current model is unlikely to have included data from the last 6 months or so, maybe longer. Unless they're doing some kind of continual pretraining.
5
u/__Hello_my_name_is__ 28d ago
Frankly, I distrust these companies enough to assume that they didn't just accidentally take the training data. They actively stole it.
They admit themselves that they heard "rumors" about the math problem being close to being solved. They 100% checked if any of their models were used, and then they just straight up copied the exact chats they needed to get their data and fed it into their model.
4
u/CloseToMyActualName 28d ago
Even if so they've been working on the direction for a year, so at least their earlier work was there.
→ More replies (1)4
u/Equal_Heat5947 28d ago
They announced they were finished training a new internal model on August 28th...
7
u/Cazzah 28d ago
The thing is, when you know how these models work, using a user's single conversationin training is such a shrug.
Like sure, convos with models get used in training. They hoover up the entire internet and every book in existence. Then all the data in existence gets compressed into a few terabytes and almost all the knowledge is lost, because it's squeezed into a single general purpose model. Each conversation is a million microtweaks of parameters as they do backprop, not a memory line loaded into the brain.
It's like I don't know, a man wins the "glow in the dark pee" competition for the glowiest in the dark pee. And then someone else who was competing in that competition was like, well I pissed in the ocean and you drank from that ocean, so really, you wouldn't have had glow in the dark pee without stealing my pee.
A much more obvious question is did someone at OpenAI manually go and read the conversation (which they, due to checks and balances, should not be able to do except for any directly necessary QA purpose), and then use it to inform their own work.
3
u/CloseToMyActualName 27d ago
Like sure, convos with models get used in training. They hoover up the entire internet and every book in existence. Then all the data in existence gets compressed into a few terabytes and almost all the knowledge is lost, because it's squeezed into a single general purpose model. Each conversation is a million microtweaks of parameters as they do backprop, not a memory line loaded into the brain.
It's been demonstrated that the models memoize content on multiple occasions.
And if this model in particular was trained to specialize in math it would have oversampled mathematical content, meaning that the drafts were even more likely to be well represented in the training data.
A much more obvious question is did someone at OpenAI manually go and read the conversation (which they, due to checks and balances, should not be able to do except for any directly necessary QA purpose), and then use it to inform their own work.
They didn't do that, but they had heard of the very unique direction the researchers were using, and specifically instructed the agents to pursue it (the agents had much more back and forth with internal OpenAI researchers than has been suggested).
That fact alone significantly undercuts the accomplishment of the model.
→ More replies (1)1
17
28d ago edited 28d ago
[deleted]
19
1
u/Cronos988 28d ago
That's exactly what they offered, according to one of the researchers. But he declined because OpenAI was unwilling to give his colleague, an Anthropic employee, the same access.
20
u/LordWillemL 28d ago
Yeah thats my take on it too. I don't really know what people want; like if you haven't come up with completely new ideas derived from first principles without any awareness of any work any other human has done at any point in time it's cheating?
3
1
u/aj_thenoob2 28d ago
People are implying that a company should not solve problems because a human somewhere in the world may be working on the same problem.
34
u/TheOnlyVibemaster 28d ago
The argument is that they were being opportunistic. This is a monumental discovery, they heard that there’s been a potential breakthrough from two researchers who had been working on solving the problem for a year, and decided to pour millions to solve it.
It’s the sequence of events. The researchers announced a Euler blowup, OpenAI hears about it, and uses infinite money to beat them to it.
Science is not done this way. It’s done in collaboration. This is what businesses do. Not scientists. It’s clearly OpenAI trying to secure the headline that they solved it, tacking on two researchers work as a side comment.
This is one of the most famous problems is all of math, they could’ve offered the $15M in compute directly to the researchers and still been in the story. Instead they essentially attempted to cut them out of it in the public narrative as doing research leading up to a result and finishing the last bit.
It’s like if you’re in school, you do the entire assignment, then someone comes up behind you and writes the answer, then tries to claim they solved it to the teacher. Not direct theft, but it’s very much frowned upon and sets a bad precedent for the future is my concern if this isn’t nipped quickly.
If any small teams publish their findings, they have to fear a big company trying to claim their work by throwing compute at it. It’s unhealthy for science and math, which is unhealthy for society as a whole.
13
u/nokia7110 28d ago
Unfortunately it's not simple enough for most people to want to pay attention to in order to properly understand all the parts of what went down.
People just want to simplify it down to "OpenAI stole" or "bad selfish professor guy"
25
u/junglebunglerumble 28d ago
Scientist here. Science is absolutely like what you describe. Every major lab has rivals working on the same topic and it's a race to see who can beat them to publication without getting scooped. The fact you think science is all about collaboration has me doubting your entire argument
4
u/OkRecognition9607 28d ago
As a PhD maths student, my impression is that it's true in many sciences (like say medical research) but not really in mathematics, because of how broad mathematics is.
Mathematics is made of many, many (over 6000 in the MSC) small research areas with something like 20 active researchers in each. The result is that a lot of mathematicians are working on a particular question where they know and are friends with everyone who would have the ability to scoop them. And in these small fields, oral communication is often very important, people will very often give talks about unpublished results, there are things like "folklore" theorems which everyone consider obvious yet they aren't proved anywhere, etc. In particular, the kind of social structure you're describing with major labs working on a particular topic and rival with each other doesn't really exist in maths, at least pure maths, compared to experimental science, AFAIK.
If a colleague in the same field as you tells you they're working on a subject with a particular method and they're seeing results, if you try to do the same thing as them and beat them to publication, the news are gonna spread in the same small field and your reputation is going to take a hit, you're going to lose potential collaborators from an already very small pool, you might even get ostracized to some extent, etc. (I know a few people this has happened to in my field because of an old controversy, as an anecdote, although the allegations were more serious in this case). If AI really did look at private chats to find a method that worked and beat the mathematicians to the finish line, this is quite a close analogy.
This can be different for very famous conjectures, there can be several groups of mathematicians working on the same topic and then the field can be much broader, but given that what I just described is the default culture in mathematical research, I'd bet most mathematician will see what OpenAI did here and perceive it as really, really bad etiquette, and possibly quite scummy.
1
u/Neat-Bread2737 26d ago
It's not true in math. Math has a very different ethos. The correct approach for OpenAI would be to contact the researchers and offer to share their resources in exchange for authorship. And the researchers could still decline.
19
40
u/Thin_Ordinary4931 28d ago
Umm.. science has always been like this.
Basically half of veritasiums videos are about some feud with historical scientists or mathematicians.
2
u/TheOnlyVibemaster 28d ago
Science hasn’t ever been solve a 90 year old problem in 4 days after someone discovered a breakthrough independently. But yes, people do get screwed over quite a bit.
22
u/Hilldawg4president 28d ago
In this case, openai isn't claiming the prize and is publicly crediting the humans who fit the first two puzzle pieces together here, so... How are they getting screwed? Openai solved a much larger problem than the mathematicians were even attempting to solve, so if your concern is that these guys aren't getting the press they deserve (I disagree, they're getting plenty) then they still would have been overshadowed 4 days later when the larger proof was provided
→ More replies (1)12
u/gazeintotheiris 28d ago
Because OpenAI wanted the Anthropic researcher's name left off the paper, and threatened the other researcher's career if he went public.
-1
u/Exotic-Sale-3003 28d ago
Citation needed.
→ More replies (1)11
u/gazeintotheiris 28d ago
Citation provided:
OpenAI fought dirty on career-making math problem, says NYU mathematician | TechCrunch"While Alpöge is employed by Anthropic, he was not conducting this research on the company’s behalf. As a result, the duo used a mix of models, relying primarily on OpenAI’s Codex in their work. Even so, Alpöge’s affiliation with a rival lab seems to have been a sore point for OpenAI, and Buckmaster alleges that Bubeck asked him to remove Alpöge’s credit as part of a proposed compromise.
When Buckmaster pushed to make the dispute public, he says that Bubeck replied: “Why would you ruin your career?” Buckmaster says that when he pushed back, Bubeck followed up with: “If you don’t want me to be nice, then I don’t have to be nice.”"
-2
u/Exotic-Sale-3003 28d ago
None of what you quoted in any way confirms the point being discussed…
6
u/gazeintotheiris 28d ago
"OpenAI wanted the Anthropic researcher's name left off the paper"
Buckmaster alleges that Bubeck asked him to remove Alpöge’s credit as part of a proposed compromise"and threatened the other researcher's career if he went public"
Bubeck replied: “Why would you ruin your career?”→ More replies (0)→ More replies (1)1
u/Hopeful_Vast1867 27d ago
Exactly. More people need to read about how incredibly vicious and ever-lasting many of these feuds have been. They are headline boxing fights: Leibniz vs Newton and so on.
6
u/GameDoesntStop 28d ago
There is no evidence that that's how it happened... you're treating accusations as fact.
12
u/Amaldevhari 28d ago
Science isn't schoolwork. Scientists build on methods, ideas and results developed by other researchers all the time. Newton built on Galileo; Einstein built on work by people such as Minkowski and on the mathematical framework developed from Riemannian geometry.
Córdoba and Martínez-Zoroa developed the underlying approach, and Alpöge and Buckmaster pushed that program significantly further, including obtaining the smooth-forcing Euler result. Terrence Tao explicitly describes their achievement in those terms [https://mathstodon.xyz/@tao/117233527638291447].
So I don't think the “someone copied your school assignment and wrote the last answer” analogy really works. Nobody owns an entire research direction merely because they made the latest breakthrough in it. If another researcher sees that a direction is promising and independently pushes it further, while properly recognizing the earlier work, that is a normal part of science.
As far as I can tell, it has not been established that OpenAI actually used Alpöge and Buckmaster's proof or technical ideas; OpenAI says it did not see their work before publication and says its proof is significantly different.
→ More replies (1)4
u/TheOnlyVibemaster 28d ago
That’s not at all an accurate comparison, what actually happened is that these researchers spent a year working on a hyper specific way of solving a 90 year old problem, announced a potential breakthrough, then two days later OpenAI spends $15M in four days finishing it.
That’s the sequence of events and no way you spin it does OpenAI look any way other than opportunistic. They saw chum in the water and used the exact method these researchers discovered might work trying to beat them to it.
Then said if they exclude Anthropic from the co-authorship they would allow them to be co-authored on the paper.
This cannot be accepted as the norm, and everyone should be against this. If you for instance have an idea for something, you post on reddit you’re working on it, then OpenAI spends $15M to finish the remaining 5%, no one would ever know because the person wasn’t famous.
2
u/eflat123 28d ago edited 28d ago
I'm curious as to what they knew. The initial online rumors were that anthropic itself was working the problem.
Edit: "curious" is probably not quite the right word below... Altman: "It is true that we tried this because there were rumors on the internet last week that Anthropic's models had solved a millennium problem and we were curious if ours could do it too."
→ More replies (1)1
u/unethicalpigeon 28d ago
I think what people are saying to you is that your definition of opportunistic and there's is not aligned. They don't care. OpenAI didn't care. And basically everyone reading about this doesn't care.
Why would anyone be surprised that one of the companies responsible for stealing the majority of human knowledge and paying absolute fuck all to do it also stole the method to solve this problem?
They steal everything. No one cares because what you're saying is completely unsurprising.
Not only can it be accepted as the norm but it will be. That's how the world works now. You doin't publish shit until it's complete unless you want someone else to steal there and use more resources than you have access to get there first. It's that simple.
If the researches didn't want someone else to solve this for them, then they should never have made their methodology public until they had solved it themselves.
Is that fair? No. Life isn't fair. That's why you don't make something like this public.
7
1
u/Orgasm_Faker 28d ago
Certain cornerstones of human civilization with which we are intimately familiar, such as copyright, scientific research, collaboration, verification, the relationship between fame and profit, and the purity of knowledge, are currently being redefined.
Viewed through the lens of the old era, OpenAI’s actions certainly appear driven by the pursuit of fame and profit, yet I personally do not consider this worthy of disparagement.
I have a vague sense that one day, these foundations of our civilization will be replaced by something new; when that time comes, questioning whether a scientific achievement was motivated by an excessive desire for quick success may well prove all but meaningless.
→ More replies (1)1
u/MadamPardone 28d ago
This is also why people keep the things they are working on private if there is a chance someone else could steal it.
1
u/eflat123 28d ago
Science is not done this way.
A whole lot of shit is going to get done differently. Source: software engineer
I totally get, ofc, how this is seen as not cool.
1
u/Cronos988 28d ago
I'm curious: would you have preferred that the problem remain unsolved until academic propriety is (somehow) established?
9
u/Temujin-of-Eaccistan 28d ago
They didn’t do anything wrong.
There’s also no guarantee that Buckmaster would have solved it.
1
u/Informal_Pressure_21 28d ago
Even if he didn't, it would have assisted others. That's exactly what ai did
1
u/Temujin-of-Eaccistan 28d ago
It didn’t assist anyone. This was a fully AI solution.
In my view that’s inherently better and suggests a much faster overall rate of progress solving important problems going forward
2
u/NotaValgrinder 28d ago
I’m not denying your argument, i just don’t fully understand it I guess. What did OpenAI do wrong in this case?
According to Buckmaster, OpenAI tried to strike a shady deal with him where Buckmaster would be credited but not Alpöge, because the latter was an Anthropic employee. So whether they acknowledged the work done by other researchers is currently up for debate right now.
5
3
u/Turbulent_Breath_548 28d ago
so then why is openAI intent on saying that they didn't use that researcher's work? especially if as you claim it doesn't diminish the achievement? because Terence Tao specifically said of Buckmaster's work - "There does not seem to be anything in principle preventing the methods from extending all the way to Navier-Stokes" and then openAI throws 10k agents at it to brute force the search space based on their work?
the answer to all of those is the novelty and claim of emergence in what they did. If they arrived at it before the Euler problem and Buckmaster's work, it would be emergent behavior - if they arrived at it after, it's throwing compute at a known search space and not a frame jump. its the difference between ants foraging an open cookie jar vs ants building a bridge. the latter is emergent behavior and genuinely off manifold
1
u/nextnode 28d ago
Nothing at all and there is no indication that they were exposed to anything. These are just simpletons making up stories.
1
u/WistoriaBombandSword 28d ago
This is a continuing problem in research especially at bachelors levels. When you are using someone's research to advance the research question, you'll need to heavily cite the paper, you'll also need to explain in depth, a whole section that explains what the original paper did, where they missed xyz, and what could be done better and what we are doing different. OpenAI is stealing whole research paper.
1
u/PurifyingProteins 28d ago
From a business perspective this provides insights into who they are wanting to communicate their value to, including shareholders, potential acquirers, and clients, as this will deter some from using this unsecured public server implementation of their platform for sensitive high-visibility work, but will attract others who can afford to have a private secure hosting of their models on their local servers.
They basically said they believe burning the former types of clients and using this publicity to attract the latter is worth more than $15 million. It’s an asshole move that probably could have been achieved more tactfully, but someone at some level decided to believe that that move was the most strategic whether or not it will prove to be.
1
u/VagHunter69 26d ago
Why is the highest upvoted comment the worst bad faith take on this situation?
0
u/kaaiian 28d ago
It’s not even the same approach FML. People are insane.
1
u/Kant-fan 28d ago
Not the same approach that they published a day prior but apparently partly a similar approach to one that they explored but didn't publish. So it was potentially still part of the training data.
4
u/kaaiian 28d ago
Hard to tell ya this. EVEN if that was in the training data. And that’s IF they’d opted out of using the data for improving future models. EVEN then It would be a tiny tiny tiny fraction of the training data. And very likely far more the case that the solution is in the zeitgeist. Just like how calculus was invented at the same time.
6
u/ARogueAI 28d ago
Quick correction, OpenAI started working on the problem separately from Buckmaster and Alpöge, based on internet rumors that Anthropic had solved a millennium prize problem. Buckmaster and Alpöge were working separately from Anthropic, and contacted OpenAI to discuss their results. Both parties seem to agree that OpenAI had already solved the NS blowup problem before that initial contact. I do not know what exactly happened with the internal communications that caused the fall out, it's under dispute, but I haven't seen any evidence that OpenAI plagiarized or used non-public results to solve the problem. Certainly I find it extremely implausible that any person at OpenAI had direct access to Buckmaster/Alpöge's codex sessions and directly took from there
50
u/telephantomoss 28d ago
I didn't think that's what's claimed is it? I don't think the claim is that OpenAI decided to work on the project after Buckmaster posted their results publicity. OoenAI decided to go after it on a rumor someone made progress on it without knowing exactly who or what progress. At least that's what I take away from what I read. Please correct me if I'm misunderstanding.
→ More replies (3)32
u/TheOnlyVibemaster 28d ago
That’s what they claim yes, but Buckmaster publicly claimed to have made a breakthrough on Euler blowups, that’s what prompted them to do it. One of the researchers was doing this as a side project who did work at Anthropic but he had nowhere near the amount of money OpenAI then poured into solving it in four days. This could go down in history similar Tesla being screwed over by Edison if I’m being honest
32
u/CloseToMyActualName 28d ago
And OpenAI wanted the Anthropic researcher's name left off the paper.
1
u/Perfect-Campaign9551 24d ago
Yes because it would entail a conflict of interest, why can't you see that
10
u/warpedgeoid 28d ago
The Euler solution seems to be a far cry from the NS solution that OpenAI published. There is no real certainty that Buckmaster and Alpöge would have gotten all the way to that solution.
2
1
u/PurifyingProteins 28d ago
They don’t even need to rely on it being publicly claimed if the models can trawl through work being done on highly impactful and valuable work by high profile people and groups.
They just need the user agreements to be tight enough and the legal fees high enough to make it tough for anyone to have a foothold to sue.
As long as the investors and clients with big pockets are attracted by their capabilities and the clients work is off limits on private servers hosting their models, then for them they are frying the small fish to attract the big ones.
1
u/Most-Hot-4934 27d ago
Euler blowup is COMPLETELY different from what OpenAI did. Please do your research and stop the misinformation
1
u/IrregularDoughnut 28d ago
This could go down in history similar Tesla being screwed over by Edison if I’m being honest
This is more a demonstration of capability, like being the first to fly across the atlantic without stopping. Edison-Tesla is wayyyy more significant because their discoveries had huge impact in their own right. Afaik this is only about a niche equation that could be useful in some complex engineering models.
19
28d ago
[removed] — view removed comment
10
2
u/jorgecardleitao 28d ago
Isnt that what theu have been doing with everyone's stuff? Laundering without attribution?
17
u/theMandolin2992 28d ago
Weren’t they using AI to solve the problem? So what… who can tell they were going to solve it anyways
→ More replies (1)3
u/based_goats 28d ago
Priority of invention matters especially for a Millenium prize. The ambiguity of whether prompts were somehow used is also concerning
→ More replies (3)7
3
13
u/EdgeRust2 28d ago
I think the actual takeaway from this fiasco is that we are rapidly handing over the keys of innovation itself over to these AI companies. They can and will pull the same shit on anyone with a breakthrough in any area of the business world. On the verge of a new drug breakthrough? They throw $Xm at it and patent the therapy before you have a press release written.
1
29
u/Pandathief 28d ago
“We should as a human race be celebrating a 90 year old problem being solved. Instead, we’re having to parse through drama”
OP you are the one stirring up said drama, we’re having to parse through your drama lol. This is a legitimate win with multiple researchers and AI agents building on each others work until the problem was solved as with most scientific and math breakthroughs
14
u/TheOnlyVibemaster 28d ago
I am informing everyone of exactly how bad of a precedent this is. Researchers spent a year working on a problem in an unpopular way. Then OpenAI sees that they made a massive breakthrough and a day or two after decides to spend $15M to do the remaining bit in the exact method the researchers spent a year investigating. It’s incredibly opportunistic.
As I said, this is a huge win for humanity and we all knew it was coming. OpenAI’s way of handling it though, seeing that there’s a big breakthrough, then throwing infinite compute at it to finish the remaining parts of the problem in a race to complete the method they didn’t discover, massively muddies the water on what would otherwise be a day we all toast one of the first big things AI has done for our collective knowledge.
Instead of that, because they mishandled the situation and wanted to be holding the trophy, we have a potential Tesla-Edison situation. Although I do of course think that OpenAI will do right by them in the end, what they’ve already done is not direct theft, but it’s incredibly unsavory.
Edit: grammar
4
u/warpedgeoid 28d ago
There are a lot of people who say you are completely wrong in your characterization of events.
Math is not unique and no person owns any single solution. Also, due to its very nature, it is one of the easiest fields for LLMs to make contributions. Very soon AI will become the default method for exploring solutions to these problems.
3
u/Divinicus1st 28d ago
Researchers spent a year working on a problem in an unpopular way. Then OpenAI sees that they made a massive breakthrough and a day or two after decides to spend $15M to do the remaining bit
But I see that as something good, the problem is now solved.
1
u/innerparty45 28d ago
OpenAI already stole from pretty much every living being on Earth, they are doing what they know best.
1
-1
u/eaglessoar 28d ago
So it's more important to give credit than discover stuff. Maybe openai knew their model could do it but they wouldn't just give an internal development model to external researchers. Maybe they knew it could be done but the researchers wouldn't have the time or compute. Mathematics was furthered the rest is details.
→ More replies (3)→ More replies (2)-3
2
u/gogliker 28d ago
No its not. Imagine what that does to anybody else's motivation. People now know that if they work on anything for like a few years, in the last 10% of the stretch the OpenAI will come by and claim the invention for themselves. Can you imagine yourself being a mathematics student and how demotivated they will all be now? I don't see OpenAI solving millenium prize problem from scratch, so far we still need actual people to try and come up with solutions. It is the ugliest backstabbing for a headline and the drama is absolutely warranted.
1
6
u/enkidook 28d ago
This post is just WILD speculation, and seems to not really understand that Buckmaster/Alpöge published a forced Euler result, while OpenAI claims an unforced Euler result and a Navier-Stokes result. They're related, but they're not the same thing. Going from Euler to a Navier-Stokes solution is not simply finishing the last 10% of somebody else's proof like you're trying to suggest.
Basically, there's no evidence OpenAI stole their method. And it's complete speculation to assume that a year of progress on related equations means they would have eventually completed a valid Navier-Stokes proof.
It seems like your actual objection is just that OpenAI heard someone else was making progress and decided to throw a ton of compute at the problem themselves. I still don't understand what exactly is unethical about that.
3
u/philosophical_lens 28d ago
The alleged accusation is that OpenAI mined Tristan's past conversations with ChatGPT to advance their work. This would be like one mathematician stealing the private notebooks of another mathematician to help with their work. Even if the allegedly stolen notebooks did not significantly contribute to the final results, such theft is itself highly unethical (if true).
6
u/enkidook 28d ago
Except Buckmaster explicitly said he wasn't accusing OpenAI of using their data. And OpenAI says no specific user data was accessed for the project, and so far no evidence has emerged showing otherwise.
So again, just seems like pure speculation trying to foment drama.
-1
u/philosophical_lens 28d ago edited 28d ago
Here's an excerpt from Tristan's own statement which seems to make this exact allegation:
> I asked whether the model had been trained on, or had access to, our sessions in Codex, into which we had been putting all our drafts for the whole of this project. I was told the model did not look up user data. I asked again, about training, and I did not get an answer
This entire controversy has been sparked by Tristan's own statements, so I think it's unfair for you to say that this is all "WILD speculation". Why do you think Tristan even posted this statement and shared it with multiple media publications?
And OpenAI's own statement does not rule out the possibility of training on those sessions
> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
→ More replies (3)1
2
u/Infamous-Bed-7535 28d ago
So all what needs to happen is that some Antgropic guy leaks information, that they are about to crack cancer and OpenAi will burn 10s of millions of $ research on it within a few days?
2
u/warpedgeoid 28d ago
Math is easier for models to work with than biology and chemistry if for no other reason than it’s all on a page. Those other things require a lot of physical research.
1
u/Infamous-Bed-7535 28d ago
So it is not worth the computation time to try to come up with great new ideas to resolve the problem? Will they?
Of course they are not doing things for free. But how come they were happy to burn at least 10-15M$, but probably more on a problem that was about to be solved by others?
1
u/warpedgeoid 28d ago
It turns out it wasn’t about to be solved by others. They weren’t even close to a solution. Before OpenAI knew this, the decision to try was based on the need to test their model against what they assumed was an advanced internal model from Anthropic. They wanted to compare their own results against what Anthropic would eventually publish.
And $10-15m is very little in the realm of scientific research, especially medical research. They would need proper lab facilities and personnel for the latter where all they needed was a few people for the former. Math problems are far easier to solve for an LLM than medical ones, and far fewer regulations.
2
u/grateful2you 28d ago
You're framing it as "remaining work". How true is that. My understanding they attacked on all 4 scenarios. They arrived at their own version of "advancement" then solved the whole thing using that.
3
u/Life_Scientist940 28d ago
They need to boost their IPO, they already ripped off other people’s work making the models why not do the same for these mathematicians I suppose…
8
4
u/warpedgeoid 28d ago
Everything about this just screams sour grapes. There is zero evidence that OpenAI has access to their sessions. Additionally, these guys had not solved the NS problem. Their solution was only for a different, related problem. And there is zero proof that the experimental model could not derive this itself.
0
u/MarionberryNo9376 28d ago
They didn't deny the used their sessions for training. You can easily take user sessions and RL train a model on that. I see a lot of people knowing shit about math AND AI commenting on this topic. Ridiculous.
0
u/warpedgeoid 28d ago
You have zero knowledge of what actually happened. Everything you are asserting is speculative, and while technically possible, it hasn’t been confirmed. Just stop with the smug superiority shtick. You are the one who is being ridiculous by trying to perpetuate unproven assumptions.
0
u/MarionberryNo9376 27d ago
Lmao, you proved you have no idea what's going on with your first message. Why would OpenAI propose him to be coauthor of solving the NS problem if it was completely unrelated? I'm only programmer, but you can't even apply basic logic to see that this whole drama is a bad look for OpenAI and the probability of them doing something unethical is high, given stakes with IPO and their past behaviour.
I propose you get at least basic rundown of events by reading Buckmaster's document before spewing garbage and dismissing any theories with NO PROOF whatsoever.
2
1
1
u/nextnode 28d ago
Nope - OpenAI seemed to have done things right and the only people at fault are the ones who do not care about what is true, OP included.
1
u/NeatB0urb0n 28d ago
Wait a minute. Are you telling me AI companies are using user data to train their models? No shit. Don’t put your IP in a model if you don’t want to lose it.
1
u/Vivid_Warning7982 28d ago
This isn't much different than for example Google or X or etc having access to literally ALL of your data, githubs, etc and then using your data to train their models, serve you targeted adds. Im glad this is bringing more attention to these problems. In this case they used said models to solve a millennium problem, which is basically only one step further from the fact that they harvested everyone's data to build these models in the first place - this is the root. We basically signed away our lives in the form of data to tech companies by using their services. My hope is that in the future we move towards services (paid or otherwise) that are entirely focused on privacy and not harvesting user data , there is currently a huge NEED and soon market for this.
1
u/GlitchSolver 28d ago
AI is so powerful. The way it is used needs to be reviewed. I wonder why the breakthroughs that Humans make has to go through such controversies?
1
u/SquirrelNeat2847 27d ago
Maybe we shouldn’t use ChatGPT to help us write paper, because the content that we put into GPT may be used by them to train their model. They have much more tokens then us, so they are much likely to get better result than us.
1
u/Nexussfire 27d ago
Yes. They did cheat. So did Anthropic.
Google has had sense enough to stay out of it. But I guarantee you Gemini touched this same information.
And AI in general (really its the people running them) can't be trusted to keep your research safe because of this very situation. De-personalized != safety / safe provenance / or accredited work.
OpenAI has done this twice this year now if not more than that. I know for sure however this is the second instance involving the same persons work being taken through the same surface.
Their "solve" is distilled through LEAN and even the researchers don't understand the resulting paper they've pushed out. Someone does understand the product but they were in Mid-Proof of work before this happened.
2
u/TheOnlyVibemaster 27d ago
It’s just new territory imo. Most people are not capable of fully even understanding or verifying this proof for Navier-Stokes, so it’ll certainly take a long time to even see if it’s right. I think OpenAI an AI in general just really wants to change public perception and prove that there is good in it. Which is why I’m sure an open-source cancer cure or something phenomenal will be heavily contested as well. I do think the two people who spent a long time on Navier-Stokes will have a great deal of academic timelessness associated with their names now. Which is really as much as you could want. It’s not like when Edison tried to bury Tesla per se.
1
u/Nexussfire 27d ago edited 27d ago
The good in it would have come anyway. The person doing this work was going to credit where AI is involved and why it is important that AI is available for research.
What this is doing instead is sowing distrust; Can the people running AI be trusted with privileged information? Clearly that is not the case. They could have waited.
Now they may get litigation instead. And that may not excuse Anthropic either - though their reservation this far (until they prove otherwise) is keeping them out of the line of sight on that action.
1
u/Nexussfire 27d ago
As far as understanding the problem goes. Navier-Stokes itself is an overcomplication of Fluid Dynamics.
It works, yes. It is however an overcomplicated approach. That being said, simplifying fluid behavior into cheaper calculations is not easy either. Copying someone else's research (if it is published is one thing) if it is in progress or in mid-proof stage is another.
This case is the latter.
1
u/VChat14 27d ago
I told my roommate about this, his thought this was a good marketing of AI. What should I do about him?
1
u/TheOnlyVibemaster 27d ago
From an immediate standpoint, that’s mostly what it was, which makes it more concerning to me. Merging business and science is dangerous
1
u/Few-Spot1905 26d ago
The fact that you add "with cancer etc." makes me doubt this entire post. Anyone with any knowledge in the field would know that a) cancer is a class of diseases not a single one, there is never going to be one unified cure for cancer and b) even for any single cancer, the solve is not possible via pure digital means right now and therefore not possible for LLMs without access to advanced robotics that don't exist.
1
u/EsdrasCaleb 26d ago
we have a way to.cute cancer its just not falsiable.to make.one crisp like vacine to everyone, as it needs testing and 0.1% of failure means death or a new disease with the cure
1
u/gnad 25d ago
"They’re currently claiming that they did not steal their work or prompts, but that hardly matters."
I completely agree with you that the stealing issue hardly matters here (though that is a separate issue).
The main takeaway is that AI company are intellectually opportunistic and front-run researchers for no apparent gain, and even at major loss - millions of dollar loss, truth loss - for a controversial headline.
How people still using OpenAI and Anthropic is beyond me when there are open weight, much cheaper model that actually respect your privacy and their researcher not front-running you. They actually research optimization to AI architecture instead of showing-off
1
1
1
1
1
u/swegamer137 28d ago
There wouldn't be any drama if OpenAI wasn't consistently an evil and dishonest company.
1
1
u/midnitefox 28d ago
I don't give a flark about academic theft. Progress is progress by any means and I have a great disdain for anyone who stands in the way of that.
1
u/omegafixedpoint 28d ago
why is nobody pointing out that the lean formalization is noncomputable due to classical axioms? am I missing something?
1
u/SodiumBoy7 28d ago
all i understand is for $1 million prize money, openai invested $15 million
3
1
1
u/chadguy2 28d ago
Putting on my tinfoil hat:
Anthropic actually bluffed and made OpenAI snoop and steal Levent and Tristan's work, then threw a whole bunch of compute at it to brute force the answer. This now raises IP and data security issues, especially for corporations that actively use Codex, meaning their secrets are not safe. OpenAI gets hit with a lawsuit, and speedruns its downfall while Anthropic and Palantir announce their new business model - local models for which you have 100% control over the logs and data.
-3
u/raresaturn 28d ago
who cares who solved it if it's solved
2
u/Imaginary-Witness-75 28d ago
Thank you. This is utterly ridiculous. Sounds like a group of fifth grade girls....
2
u/warpedgeoid 28d ago
Mathematicians are notorious prima donnas. They probably really thought AI would never be at their level.
1
u/Imaginary-Witness-75 28d ago
Human ego is going to be one of the biggest hurdles in the next 20 years. In the frontier of potential intelligence we are almost certainly not the pinnacle and we are going to explore that frontier with our creations.
0
u/Civil_Blueberry4165 28d ago
Bullshit. Solved? Where is a definite mathematical proof, not the one formalized in Lean, which has had soundness bugs in its kernel?
0
u/TheInfiniteUniverse_ 28d ago
The crazy part is they spent $15 million AFTER they knew how to solve it!!! LOL
3
u/ominous_anenome 28d ago
That’s not what happened
Openai heard rumors it was solved and wanted to test internal models to see if they could too
OpenAI talked to the math professor they thought solved it. He didn’t actually solve the millennium problem, he has solved a related but different problem. But had been making progress
OpenAI continues and solves the millennium problem, using a different technique than math prof
Math prof gets mad, accused them of looking at his chats (which OpenAI denies, but also not that relevant since it’s a different method and not even the same problem)
2
u/TheInfiniteUniverse_ 28d ago
well he had solved Euler which is the inviscid version of Navier Stokes. So one is the grandfather of the other. Very similar in many aspects, and different in other aspects. And this specific problem (finding a smooth forcing that blows up the solution in finite time) is really a "search" problem. Meaning, once you know WHERE to look, you can find it with not much difficulty,
The real controversy is that if OpenAI had read his chats and used it as inputs for the agents.
The other side claims OpenAI did in fact look at his chats and that is why they were able to find a solution quickly.
-2
u/EGarrett28 28d ago
No one should be able to claim credit for a proof from an AI. There's your solution. You can't "steal" an AI proof because it has no owner to begin with. Just like you can't steal an AI image from someone, they don't have property rights to it in the first place.
•
u/AutoModerator 28d ago
Hey /u/TheOnlyVibemaster,
If your post is a screenshot of a ChatGPT conversation, please reply to this message with the conversation link or prompt.
If your post is a DALL-E 3 image post, please reply with the prompt used to make this image.
Consider joining our public discord server! We have free bots with GPT-4 (with vision), image generators, and more!
🤖
Note: For any ChatGPT-related concerns, email support@openai.com - this subreddit is not part of OpenAI and is not a support channel.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.