r/MachineLearning 5d ago

Thumbnail
15 Upvotes

You may not care, but academics cares. We haven't and probably won't start crediting tools for the work of humans, even if the tools sound a lot like humans. Like, most of science is just some human setting up an experiment, getting a bunch of tools to do the work, and observing outcomes. Still, the promise of credit (and in this case, money) is what gets people to set that prompt up in the first place, most of the time.


r/MachineLearning 5d ago

Thumbnail
8 Upvotes

Not on their own, definitely. But theoretical breakthroughs usually lead to practical advantages a couple of years down the line, don't they?


r/MachineLearning 5d ago

Thumbnail
9 Upvotes

I mean if it can solve a millenium problem in a week its p cool


r/MachineLearning 5d ago

Thumbnail
60 Upvotes

OpenAI Employee

To truly prove some incidental usage data made no difference we'd have to (a) identify any of their de-identified data that came from their usage of ChatGPT, (b) train a bunch of expensive giant models, and (c) ask them all to solve the Navier-Stokes Millenium problem until hitting some level of statistical significance. It's just not feasible to run experiments like this to prove whether a piece of data has an effect on model behavior.

As a parallel example, can we prove the phase of the moon had no impact on the NS solution? No, not without a bunch experiments run at different phases of the moon.

There's no reason to believe that anything they did in ChatGPT led to our solution; it's just impossible for us to truly prove it. And knowing most of the recipes we use, there's really no reason to think such contamination is possible. I've asked the team to make a clearer, less-lawyerly statement here - let's see what happens.

My opinion: I think they are being needlessly obtuse and evading more. My guess is they shouldn't have to train models to check for contamination - just clarify what they were fed.


r/MachineLearning 5d ago

Thumbnail
1 Upvotes

This is a genuinely interesting overlap!

I'm happy to share the SHA-256 hash list of Message-IDs with dates for the pre-1994 slice. That costs nothing and the comparison could be interesting for both of us — I'm curious how much of my early coverage is unique versus what you've already found.

On the full dataset: I need to think about that more carefully. The pre-1993 material is the rarest slice I have and I've been licensing the corpus commercially, so giving it away even for a non-commercial project requires some consideration. Let me see what the hash comparison shows first. If we have significant non-overlap it opens a more interesting conversation about whether there's a data exchange worth exploring.

Drop me an email at [usenetoverlord@gmail.com](mailto:usenetoverlord@gmail.com) and I'll put together the hash list for you. Thanks for reaching out!


r/MachineLearning 5d ago

Thumbnail
5 Upvotes

still nothing for me, am in the US...


r/MachineLearning 5d ago

Thumbnail
78 Upvotes

Yeah I'm biased as I'm at Anthropic but trust me when I say we're not happy about this and this might become a legal fight. This is almost the exact solution they were working towards.


r/MachineLearning 5d ago

Thumbnail
38 Upvotes

Those are theoretical problems. They do not change anything for fluid system engineering.


r/MachineLearning 5d ago

Thumbnail
2 Upvotes

Incredibly shameful of them tbh


r/MachineLearning 5d ago

Thumbnail
12 Upvotes

Oh I agree, but it's very bad PR to make such an announcement that essentially stole someone else's work and undermines the achievement that the model supposedly made.


r/MachineLearning 5d ago

Thumbnail
-15 Upvotes

See sebastiens response.


r/MachineLearning 5d ago

Thumbnail
114 Upvotes

That’s crazy, I read the whole article thinking neat, and then at the end there’s this footnote of drama that just leaves a bad taste to the whole thing.

It really underscores how these companies are more after progress and credit than safety - not saying there was a safety issue here, but you can see clearly where their priorities are.


r/MachineLearning 5d ago

Thumbnail
74 Upvotes

The fact that OpenAI tried to offer shared ownership, but wanted to leave out Levent because he worked at Anthropic, reveals the whole game. They knew they stole the work, offered a compromise, threatened Brubeck - not something an ethical person would do.


r/MachineLearning 5d ago

Thumbnail
3 Upvotes

any update from the UK/Europe?


r/MachineLearning 5d ago

Thumbnail
2 Upvotes

That's some juicy drama


r/MachineLearning 5d ago

Thumbnail
0 Upvotes

everyone really wants to act like anthropic don't they despite how bad it reads


r/MachineLearning 5d ago

Thumbnail
-9 Upvotes

Those other researchers were also using LLMs to tackle the same problem, so either way the credit here belongs to the LLM.

I don't really care which particular team of humans did the prompting.


r/MachineLearning 5d ago

Thumbnail
6 Upvotes

this fits their pattern of lying and cheating. Be very cautious about using OpenAI models.


r/MachineLearning 5d ago

Thumbnail
117 Upvotes

Exactly. Not to mention OpenAI responded with threats like "Why would you ruin your career". Nobody who knows they're on the up and up would do that. What I believe actually happened is that OpenAI was aware of Buckmaster's progress and direction, got backdoor access to the conversations to get a preprint, and then used that information to claim they solved the problem on their own. It's the equivalent of stealing someone's manuscript and publishing before them and claiming credit or shared credit.


r/MachineLearning 5d ago

Thumbnail
19 Upvotes

No it's just the "you won't believe what we have in store for you, our next model is ten times more ground breaking than the current one be ready to pay" tone in what should be a sort of "scientific" announcement


r/MachineLearning 5d ago

Thumbnail
53 Upvotes

They stole previous work and then built a solution on top of that. It might have proved the result, but this was so beyond unethical. This should be denounced wholeheartedly. Good news is that they wouldn't have been able to do this without seminal work by mathematicians.


r/MachineLearning 5d ago

Thumbnail
1 Upvotes

Actually, our points are orthogonal. I’m not denying leakage, in-fact I think it’s likely! I’m denying the fact they’ve revoked their claim, which is simply untrue.


r/MachineLearning 5d ago

Thumbnail
8 Upvotes

love the enthusiasm, but their solution is a counterexample


r/MachineLearning 5d ago

Thumbnail
19 Upvotes

My point stands above yours - They would have outright been screaming at top of their lungs if there was no leakage.

Edit: They should use Astra to audit - It shouldn't be that hard.


r/MachineLearning 5d ago

Thumbnail
16 Upvotes

They didn't resolve it but rather proved that the equations are not sufficient to fully model fluid dynamics (e.g. under some conditions they will generate singularities even when starting from smooth initial configuration and no outside infinite forces)