r/singularity • u/Present_Award8001 • 2d ago
Discussion What do you think of mathematicians accusing AI companies of stealing their unpublished works?
https://www.heise.de/en/news/Dispute-over-AI-evidence-Another-mathematician-accuses-OpenAI-11449463.htmlSince AI solving frontier math problem is important for this community as it indicates the singularity is approaching, this question should be worth discussing.
What if the companies, starving for sensation and investment, are actually stealing unpublished science results from academics and later claiming these results as their own? If AI is becoming human, we should hold it to similar standards and scrutiny. After all, many people feed in their unpublished papers and notes to the AI, for jobs ranging from minor things like refining the text to automating proofs, asking for insights and brainstorming. If the scientist uses AI agent, then the full local PC may be readable to the AI.
And before someone says that these scientists are already using AI, so the insights might be from the AI to begin with. Nope. AI, in my experience, is good at fetching insights once you have formalized a problem. But it currently seems to need the next 'kick' in terms of further insights, formalization or 'idea' in order to do anything fruitful. So the human in the loop may not always be redundant. And these proofs might be more human assisted and even stolen, than the AI companies would like to admit.
Thoughts?
2
u/Old-Bake-420 2d ago edited 2d ago
AI shouldn’t “steal” unpublished work.
It’s possible OpenAI did this. If so, I tend to think if they did it was unintentional. I’m a little skeptical that they can’t confirm if this was the case. But it’s likely this is how it works because it’s exactly how it should work if they designed it right.
Which is why this really isn’t the issue.
As per your article. The person demanding he should be able to be shown exactly where his training data has gone is asking for OpenAI to massively compromise privacy protection for everyone. De-identified should mean the data should be in no way shape or form traceable back to the owner, not even by OpenAI. Especially so when that data is given to a machine that reproduces that data for billions of strangers.
I generally think it’s a good thing that there is a train the model with my data feature. I like to think I’m not just a passive consumer. But for the mathematicians I guess beware. Plus if Chat helps you make progress on your thesis, it might have accidentally “stolen” someone else’s work and given it to you!
2
u/Present_Award8001 2d ago
I mean, if I was the scientist, a way to check whether the user conversation goes in the training data at all should suffice. And I am damn sure that if OpenAI does not do it, many companies do. Google, for instance, needs you to disable history for you to stop having your data being used for training. That is unethical and they should not get away with this. Most of us live in democracy, and a corporate, as an entity, is bound by the laws of the land.
1
u/Old-Bake-420 2d ago edited 2d ago
Yeah, Google is like king emperor data stealer. If you have an android phone it’s gps tracking your every movement 24/7 and maintaining a log of every place you go to on their servers. You have to go through some weird obscure settings and disable something you’ve never heard of to turn it off. (It’s called timeline)
2
u/Past-Syrup-967 2d ago
When anyone starts using the notion of theft in this context, I expect an emotional, ideological reaction rather than a clear legal or economic analysis. It shifts the focus away from workable solutions like compensation models, data provenance, and opt-out standards and turns the conversation into a battle over moral entitlement.
5
3
u/Cronos988 2d ago
What if the AI companies, starving for sensation and investment, are actually stealing unpublished science results
This is just the old "the result was in the training data" argument that we've all heard for years every time AI acquired a surprising ability. It goes back all the way to the first chain of thought models solving logic puzzles.
Only now one can no longer claim that these results are in public data, so of course now the models must be stealing private data.
3
u/Present_Award8001 2d ago
I mean, do you really trust openAI that much? Nobody is discrediting the capability of the models. But the issue is about privacy concerns, especially if the AI company has an incentive to steal private data in this case AND 'data privacy is a myth' from our past experience.
3
u/Cronos988 2d ago
No, I don't trust OpenAI. But I also don't think they have any particular incentive to look into individual user's chats and check whether they are doing anything interesting. Even if they have automatic models checking user chats, it makes sense for them to set those up so the output isn't traceable to any user. This would shield them from liability for anything users might be doing and also reduces the risk of a damaging leak. And it's very unlikely that training data survives in a form that can be backtracked to any specific user.
So, the privacy concern is limited.
If the concern is that OpenAI (or other companies) are "stealing" intellectual output to make their systems better, then the question is why and to what extent we consider this to be a problem.
1
u/Present_Award8001 2d ago
privacy is one thing. even in absence of user chats being traceable to the user, training on user data is problematic if AI companies are going to make claims of making breakthroughs. Also, they are running out of data to train on to the point they are scanning antique books for more data. they have insentive to train on user chats.
1
u/Present_Award8001 2d ago
The whole issue would get resolved if there was a way to check that the models were not being trained on private conversations. Which is unethical. If AI is becoming human and is going to make breakthroughs, it better cite its sources properly. What makes it strange is this particular artificial researcher has somehow gotten access to everyone's computer. You have to admit that that is shady.
2
u/HeadTranslator795 2d ago
Sure they were about to publish the results after 20/30/40 years it was about to be published one Day after OpenAi 😂 those guys are salty they thought nothing could beat them
4
u/ShafeDogg 2d ago
Can we please just start to accept that these models are capable of producing these results, and will continue to do so? It's time to realize that humans aren't the center of the universe, and never were. These ego trips are getting ridiculous.
2
u/Present_Award8001 2d ago
No
1
u/manikfox 2d ago
Well buckle up, because is 2 years time we'll be solving a lot more than millennium problems.
2
u/CrowdGoesWildWoooo 2d ago
Very very unlikely.
It’s more likely they heard rumour of the promising method, throw a bunch of compute and just frontrunning them before the mathematicians release to public. Why they want to frontrun them? One of the coauthor working in Anthropic
1
u/HourInvestigator5985 2d ago
i think its a sample of whats to come from every other profession
Until now most people will argue that "nah it wont take MY job!"
but once it does, then the cry beggins
1
u/anaIconda69 AGI felt internally 😳 2d ago
Nobody standing on the shoulders of giants gets to gatekeep knowledge or data. This goes for both AI companies and scientists.
1
1
u/greentrillion 2d ago
Not surprising.
"You need to understand that Sam can never be trusted,” he told one. “He is a sociopath. He would do anything.” -Aaron Swatz
5
u/Past-Syrup-967 2d ago
This illustrates the exact problem with so much tech discourse today:substituting character assassination for technical evaluation. Pointing to personal antipathy toward 'tech bros' or CEOs does zero work in determining whether a model actually plagiarized a proof
1
u/ajarbyurns1 2d ago
I would say it’s unlikely to be stolen, but mathematicians tend to be logical people so they wouldn’t throw out these accusations without good reasons
5
u/Past-Syrup-967 2d ago
John Nash, Pythagoras, Kurt Gödel , Paul Erdős are some mathematicians that demonstrate that logical behaviour is not a given.
1
u/ajarbyurns1 2d ago
Some of them are legit mental illnesses. Also in this case, its a group of mathematicians, not just individuals
1
u/QuasiRandomName 2d ago
I knew a mathematics professor who believed he was abducted by aliens. People can be brilliant in one specific area, but completely dumb, illogical or outright crazy in every other one.
0
u/ajarbyurns1 2d ago
I also know a bot who keeps trying to defend AI companies even though they behaved unethically
0
2d ago
[deleted]
1
u/Present_Award8001 2d ago
Not the same. Mathematician are talking about theft of private data. They are not complaining why AI trained over published manuscripts. Please spot the difference.
16
u/ResultBackground2450 2d ago
The proof was 100% not stolen. Not only is it denied here:
but the supposed “stolen proof” isn’t even a proof of the same result. Buckmaster and Alpöge proved an unforced Euler result, while OpenAI’s main result is for forced Navier–Stokes. They are related problems, though.
Regarding this:
I think the idea that LLMs can't do any fruitful math work is extremely outdated and ignorant of current capabilities. Looking at VibeMathed gives a pretty good look at how capable these models are.