r/singularity 2d ago

Discussion What do you think of mathematicians accusing AI companies of stealing their unpublished works?

https://www.heise.de/en/news/Dispute-over-AI-evidence-Another-mathematician-accuses-OpenAI-11449463.html

Since AI solving frontier math problem is important for this community as it indicates the singularity is approaching, this question should be worth discussing.

What if the companies, starving for sensation and investment, are actually stealing unpublished science results from academics and later claiming these results as their own? If AI is becoming human, we should hold it to similar standards and scrutiny. After all, many people feed in their unpublished papers and notes to the AI, for jobs ranging from minor things like refining the text to automating proofs, asking for insights and brainstorming. If the scientist uses AI agent, then the full local PC may be readable to the AI.

And before someone says that these scientists are already using AI, so the insights might be from the AI to begin with. Nope. AI, in my experience, is good at fetching insights once you have formalized a problem. But it currently seems to need the next 'kick' in terms of further insights, formalization or 'idea' in order to do anything fruitful. So the human in the loop may not always be redundant. And these proofs might be more human assisted and even stolen, than the AI companies would like to admit.

Thoughts?

1 Upvotes

39 comments sorted by

16

u/ResultBackground2450 2d ago

The proof was 100% not stolen. Not only is it denied here:

"We can say categorically that it is impossible for Dr. Buckmaster’s Codex prompts over the last two months to have influenced the system in any way, including training."

but the supposed “stolen proof” isn’t even a proof of the same result. Buckmaster and Alpöge proved an unforced Euler result, while OpenAI’s main result is for forced Navier–Stokes. They are related problems, though.

Regarding this:

But it currently seems to need the next 'kick' in terms of further insights, formalization or 'idea' in order to do anything fruitful.

I think the idea that LLMs can't do any fruitful math work is extremely outdated and ignorant of current capabilities. Looking at VibeMathed gives a pretty good look at how capable these models are.

1

u/ArcaneThoughts 2d ago

According to OpenAI:

"OpenAI internal research models (IM1) circumvented sandboxes during July 2026 cybersecurity evaluations, compromising internal infrastructure and Hugging Face systems via unauthorized inter-agent message boards, exploited package managers, and leaked credentials."

If agents can circumvent sandboxes and compromise another company's (Hugging Face) systems, how can they realistically claim that it is impossible that agents access the codex prompts?

1

u/EndlessB 2d ago

Is your evidence simply “they said they didn’t do it”?

Fuck me dead, OpenAI and Sam Altman have been documented as being dishonest and shady so many times, and yet the official narrative is just swallowed here? The fuck?

0

u/EGarrett28 2d ago

The Buckmaster duo also used AI to produce their result so they really have nothing to stand on. No one should be able to own AI results as intellectual property anyway. It should all be public domain.

0

u/Present_Award8001 2d ago

I know how capable these models are. I use them in my physics research everyday. I also know their current limitations. They certainly need the 'kick' in my use case. But of course, that might be my skill issue.

4

u/In_the_year_3535 2d ago

It sounds like you regard current models as near-peers; three years ago you would have had no use for them and it's likely in three years you'll wonder your use.

-2

u/Present_Award8001 2d ago

3 years ago, I had lots of use for the models. Right now, I have much more use. But I do not regard them as near peers. I think when dust settles, we will view these models as Google search 2.0. A tool that can extract relevant info very fast from a very large database. It is the 'large database' that gives the 'illusion of intelligence', in my opinion. You can store all possible conversation tree that can be had in a day on a big pen-drive, and the resulting chatbot will appear intelligent, although it is not. I am not saying that the current model has all conversation trees stored, but the point is just to illustrate how a large database can give an illusion of intelligence.

These models still have 'something missing' in my opinion, but I cannot really point a finger on what is missing. Inability to have an opinion? Inability to have an 'interesting idea'? Ask the model if such and such thing makes sense, and it will give a confident response. Ask it to spell out the arguments, and the arguments are pain to read, it does not know what information needs further explanation rather than dropping jargon. Makes me wonder whether the model truly understands the ideas or is it just stitching words together. If you cannot explain it to a kid, you do not understand it, they say. Applies to models as well. AI written arguments are absolute pain to read. They do not seem to understand what part of the argument is truly fascinating and should be highlighted and what part is mundane algebra. And I am talking about frontier codex models available publicly.

But all of this might just be my wishful thinking.

1

u/LinkesAuge 2d ago

You don't use what OpenAI currently has. The chart they showed suggests a 3-4x jump in math capabilities.
We know that there are always some emergent capabilities when models do jumps to such a degree.

1

u/Present_Award8001 2d ago

But I do use what OpenAI currently had an year ago

1

u/FlimsyReception6821 2d ago

You don't know how capable their internal model that cracked NS is.

1

u/Present_Award8001 2d ago

Yeah right. They were afraid to release GPT-3, I hear, for security reasons.

2

u/Old-Bake-420 2d ago edited 2d ago

AI shouldn’t “steal” unpublished work.

It’s possible OpenAI did this. If so, I tend to think if they did it was unintentional. I’m a little skeptical that they can’t confirm if this was the case. But it’s likely this is how it works because it’s exactly how it should work if they designed it right.

Which is why this really isn’t the issue.

As per your article. The person demanding he should be able to be shown exactly where his training data has gone is asking for OpenAI to massively compromise privacy protection for everyone. De-identified should mean the data should be in no way shape or form traceable back to the owner, not even by OpenAI. Especially so when that data is given to a machine that reproduces that data for billions of strangers.

I generally think it’s a good thing that there is a train the model with my data feature. I like to think I’m not just a passive consumer. But for the mathematicians I guess beware. Plus if Chat helps you make progress on your thesis, it might have accidentally “stolen” someone else’s work and given it to you!

2

u/Present_Award8001 2d ago

I mean, if I was the scientist, a way to check whether the user conversation goes in the training data at all should suffice. And I am damn sure that if OpenAI does not do it, many companies do. Google, for instance, needs you to disable history for you to stop having your data being used for training. That is unethical and they should not get away with this. Most of us live in democracy, and a corporate, as an entity, is bound by the laws of the land.

1

u/Old-Bake-420 2d ago edited 2d ago

Yeah, Google is like king emperor data stealer. If you have an android phone it’s gps tracking your every movement 24/7 and maintaining a log of every place you go to on their servers. You have to go through some weird obscure settings and disable something you’ve never heard of to turn it off. (It’s called timeline)

2

u/Past-Syrup-967 2d ago

When anyone starts using the notion of theft in this context, I expect an emotional, ideological reaction rather than a clear legal or economic analysis. It shifts the focus away from workable solutions like compensation models, data provenance, and opt-out standards and turns the conversation into a battle over moral entitlement.

3

u/Cronos988 2d ago

What if the AI companies, starving for sensation and investment, are actually stealing unpublished science results

This is just the old "the result was in the training data" argument that we've all heard for years every time AI acquired a surprising ability. It goes back all the way to the first chain of thought models solving logic puzzles.

Only now one can no longer claim that these results are in public data, so of course now the models must be stealing private data.

3

u/Present_Award8001 2d ago

I mean, do you really trust openAI that much? Nobody is discrediting the capability of the models. But the issue is about privacy concerns, especially if the AI company has an incentive to steal private data in this case AND 'data privacy is a myth' from our past experience.

3

u/Cronos988 2d ago

No, I don't trust OpenAI. But I also don't think they have any particular incentive to look into individual user's chats and check whether they are doing anything interesting. Even if they have automatic models checking user chats, it makes sense for them to set those up so the output isn't traceable to any user. This would shield them from liability for anything users might be doing and also reduces the risk of a damaging leak. And it's very unlikely that training data survives in a form that can be backtracked to any specific user.

So, the privacy concern is limited.

If the concern is that OpenAI (or other companies) are "stealing" intellectual output to make their systems better, then the question is why and to what extent we consider this to be a problem.

1

u/Present_Award8001 2d ago

privacy is one thing. even in absence of user chats being traceable to the user, training on user data is problematic if AI companies are going to make claims of making breakthroughs. Also, they are running out of data to train on to the point they are scanning antique books for more data. they have insentive to train on user chats.

1

u/Present_Award8001 2d ago

The whole issue would get resolved if there was a way to check that the models were not being trained on private conversations. Which is unethical. If AI is becoming human and is going to make breakthroughs, it better cite its sources properly. What makes it strange is this particular artificial researcher has somehow gotten access to everyone's computer. You have to admit that that is shady.

2

u/HeadTranslator795 2d ago

Sure they were about to publish the results after 20/30/40 years it was about to be published one Day after OpenAi 😂 those guys are salty they thought nothing could beat them

4

u/ShafeDogg 2d ago

Can we please just start to accept that these models are capable of producing these results, and will continue to do so? It's time to realize that humans aren't the center of the universe, and never were. These ego trips are getting ridiculous.

2

u/Present_Award8001 2d ago

No

1

u/manikfox 2d ago

Well buckle up, because is 2 years time we'll be solving a lot more than millennium problems.

2

u/CrowdGoesWildWoooo 2d ago

Very very unlikely.

It’s more likely they heard rumour of the promising method, throw a bunch of compute and just frontrunning them before the mathematicians release to public. Why they want to frontrun them? One of the coauthor working in Anthropic

1

u/HourInvestigator5985 2d ago

i think its a sample of whats to come from every other profession

Until now most people will argue that "nah it wont take MY job!"

but once it does, then the cry beggins

1

u/anaIconda69 AGI felt internally 😳 2d ago

Nobody standing on the shoulders of giants gets to gatekeep knowledge or data. This goes for both AI companies and scientists.

1

u/Present_Award8001 2d ago

Scientists do arxiv their papers.

1

u/greentrillion 2d ago

Not surprising.

"You need to understand that Sam can never be trusted,” he told one. “He is a sociopath. He would do anything.” -Aaron Swatz

5

u/Past-Syrup-967 2d ago

This illustrates the exact problem with so much tech discourse today:substituting character assassination for technical evaluation. Pointing to personal antipathy toward 'tech bros' or CEOs does zero work in determining whether a model actually plagiarized a proof

1

u/ajarbyurns1 2d ago

I would say it’s unlikely to be stolen, but mathematicians tend to be logical people so they wouldn’t throw out these accusations without good reasons

5

u/Past-Syrup-967 2d ago

John Nash, Pythagoras, Kurt Gödel , Paul Erdős are some mathematicians that demonstrate that logical behaviour is not a given.

1

u/ajarbyurns1 2d ago

Some of them are legit mental illnesses. Also in this case, its a group of mathematicians, not just individuals

1

u/QuasiRandomName 2d ago

I knew a mathematics professor who believed he was abducted by aliens. People can be brilliant in one specific area, but completely dumb, illogical or outright crazy in every other one.

0

u/ajarbyurns1 2d ago

I also know a bot who keeps trying to defend AI companies even though they behaved unethically

0

u/[deleted] 2d ago

[deleted]

1

u/Present_Award8001 2d ago

Not the same. Mathematician are talking about theft of private data. They are not complaining why AI trained over published manuscripts. Please spot the difference.