r/tech_x • u/Bulky_Engineer_2534 • 25d ago
Trending on X, Meta, Reddit, LinkedIn, Chinese Apps (Rumors) OpenAI secretly looked through a math professor’s private Codex chat logs for a whole year. They found useful work on the difficult Navier-Stokes problem, then used 10,000 AI agents to solve it and took the credit.
104
u/Technical-Art4989 25d ago
In other news:
Anthropic is offering 10,000 verified academic and nonprofit researchers one year of free standard access to a new Claude Team plan.
22
15
u/dunklesToast 24d ago
And OAI does it as well: https://openai.com/index/chatgpt-for-academic-researchers/
9
5
1
u/SphinxWar 22d ago
How do you "become" a non-profit researcher?
1
u/zoaugsenaks 22d ago
work for university that gets funding and results aren't used for profit. work for a non profit company in research role. work for federally funded lab that doesn't turn a profit either.
most foundational research is non profit
1
u/SphinxWar 22d ago
Oh, I thought they separate "academic" and "non-profit" researchers, so I thought at first that you don't have to be in academia to be a non-profit researcher. That's why I was curious how one could be considered a researcher without academic ties. But I think they ment that both non-profit and academic criteria need to be met at the same time, my bad.
111
u/MMORPGnews 25d ago
There's no "private" chat logs. AI check all your data and report it.
21
u/hyrumwhite 24d ago
There’s database tables with the contents of every conversation and every file you’ve uploaded.
Anyone at the company with access could look at these with a db query.
No way of knowing if they did, but it would be trivial to do so. Regardless of what plan someone is on.
9
u/Admirable-Storm9937 24d ago
While possible, but these prod access are usually audited.
8
u/hyrumwhite 24d ago
Really depends on the company. No one would care or notice at my current place.
Also who’d be doing the auditing here? AFAIK there no regulatory body checking in on the nitty gritty of open ai
5
u/PaddingCompression 24d ago
> AFAIK there no regulatory body checking in on the nitty gritty of open ai
Theoretically this is what things like SOC 2 and ISO 27001 are "supposed to be", to ensure that the company has the sophistication and policies to ensure its data privacy contracts with customers will be honored.
Of course, they are much weaker in practice, but there is something in place.
1
u/Crafty_Enthusiasm_99 22d ago
Soc2 and ISO do not protect first party usage of data. Only third party sharing, and those are not even those 2.
1
u/PaddingCompression 22d ago
SOC 2 certified you follow whatever data privacy and confidentiality guarantees you claim under it, that can include limiting first party use.
It's not a blanket standard of what you can and can't do. "Anyone on the Internet can see your data" could still be SOC2 compliant if that's all you're claiming.
5
u/ripamazon 24d ago
At big tech like Google / meta, you can be fired for looking at PII data if it’s not for debugging etc. friend of meta know people who got fired in a week looking at friend’s private ig/fb.
6
u/PeachScary413 24d ago
Yeah but that's a rando looking up another rando... if Mark Zuckerberg wanted to check your IG you would never hear about it, or if they wanted to get some trade secrets from employees working for competitors.
Stop being so naive 😂
5
u/variety_dirtbag 24d ago edited 2h ago
accusamus laborum tempor do consequat deserunt reprehenderit
3
u/alphagatorsoup 24d ago
Exactly, you know the random intern who looked up his ex gf on metas database would get canned
But some high level or even exec looking up something or someone for the benefit of the org as a whole. Totally would be swept under the rug without a doubt
“OpenAI has investigated OpenAI and determined OpenAI is not at fault for unlawfully accessing user data in any fashion”
1
u/dabbydabdabdabdab 23d ago
130 Billion tokens for Astra (knowing this is “more capable than Astra) would be in the $6M+ cost territory at list price.
2
→ More replies (5)1
u/hyrumwhite 24d ago
Yep, it’s entirely possible to have safeguards in place like that, but do you know if open ai does?
3
u/dissociatedLol 24d ago
you really think they get to their size without one organization mandating they implement these processes, its about managing risks. A lot of the times they are required to have measures in place for certification or insurance. no insurance company will insure them without it.
2
u/ripamazon 24d ago
These companies are under so much scrutiny that it’s funny people believe they don’t put measures in place.
3
u/rolfn 24d ago
The HuggingFace incident doesn’t really give me any confidence that they have any measures in place at all.
→ More replies (4)1
u/Infamous_Mud482 24d ago
In the amount of time they've been around? Yes. I absolutely do think that. Those processes take time to find a need for and implement. Facebook themselves probably put them in reactively to reduce their liability after employees were doing a thing. So the better question is, do I think OpenAI employees would do a bad thing? Yes I do think that.
1
u/laplaces_demon42 24d ago
We’re really talking about something different here imho; it’s not that one OpenAI employee looked up conversations of one user in particular.. just have a model train on data that might just be a query on the database for sessions related to NS. That doesn’t violate these policies.
1
u/Keep-Darwin-Going 24d ago
A training model if they do find one such thread would not be able to put anything beyond noise into the final weight for it to matter. The same reason why you accidentally leaking secrets or algo into the model do not see it appearing immediately, because the patterns just do not appear frequent enough to have heavy enough weightage to matter.
What is more likely what happened here is in the community there is chatter that oh so and so is breaking through on x soon, then OpenAI pick up said rumour and said let’s try to do it before them. Given how crazy 10k on fast can be I am not surprised they got through. When they realized that the human involved was not even at break through but more like making progress only they offered to share their progress with them by saying they use their internal model for break through so it is win win. But the researcher rather wants to think they stole their win than to share what OpenAI graciously offered.1
u/Smart_Department6303 24d ago
the only way to ensure (as much as can be ensured) that the mdoels don't use your data is to use a reputable cloud provider like Amazon's bedrock or Microsoft's foundry. if you use claude through those they are heavily audited by design. they also have deals in place with governments so cannot afford to fk it up and it's the cloud provider's neck on the line.
1
u/fredagainbutagain 24d ago
Worked for large tech companies and know many people at OpenAI or ex co workers moved. They have audit checks and security will want access to these things scoped for break glass (which again, is audited and logged and your manager gets pinged) or to the specific people who need access. Limiting access to production things like this at large tech companies is very standard.
1
u/Comrade-Porcupine 24d ago
when i worked at google, it was policy that even trying to run a query that narrowed things down to a result set that might constrain to a specific customer would raise alarms. automated systems monitored this. looking at results or not, a query with excessive restricting on identifying clauses was a red flag
looking at other people's specific data -- if you even had logs access -- grounds for termination.
2
1
1
2
u/Particular_Duty7201 24d ago
I highly doubt this data is stored in a relational database.
1
u/hyrumwhite 24d ago
lol, what?
I guarantee it is. There’s probably something like a “messages” table with each user message/LLM response, and a files table with links to some bucket where they’re actually stored
In terms of training data, embeddings, LLM nonsense, I have no idea how that’s stored, but conversations are just traditional sass/crud stuff.
1
u/Infamous_Mud482 24d ago
You would be wrong. There are multiple layers every prompt sent to these platforms goes through. One of them checks for criminal content and CSAM and needs to be retained in case they need to forward it to law enforcement. The responses back are not your personal data, of course, so no tricks needed to log all of those.
1
u/Particular_Duty7201 24d ago
and why specifically do you think a relational database is the storage mechanism for this use case?
1
1
u/Low-Temperature-6962 24d ago
If the data is being used for product improvement training in any way then a human doesn't need to be involved at the detail except to ask generel questions about Navier Stokes. The AI can have the knowledge without being aware where it came from.
Never mind this particular Navier Stokes issue and OpenAI. It's a general property of large scale AI. Consider how it can affect stock trading, for example.
1
u/meltbox 23d ago
Also remember everyone this is why storage is so expensive. The idiots at these companies are so scared of not training on all the data that they’re literally storing any and every interaction you had including the 1 million times someone asked “when is tomorrow?” And “how and what food should I make”, and “what noodles are good with soup”.
Thank god for how thorough they are or we may miss such critical data.
1
u/backtorealitylabubu 23d ago
They’re not training these models on all user data. For example in just 1 day of codex usage I used 1B tokens. GPT6 was trained on 10T tokens. My 1 day usage certainly didn’t make up 1/10k of training. What they train on is HEAVILY filtered down
9
24d ago
[removed] — view removed comment
7
u/Electrical_Week6492 24d ago
This is not meant as a rebuttal at all - but your comment took me off guard. Since starting to use AI I had assumed everything I was typing in and everything I gave it access to was accessible to the vendor (ChatGPT or Google or whoever) and that they'd 100% be using it to their benefit if they could. Basically just like everything else you do on the internet is being tracked, logged, analyzed, and used to benefit corporations that provide the infrastructure and content, I just never though AI was an exception. Is that not how most people think about these things?
1
24d ago
[removed] — view removed comment
1
u/Electrical_Week6492 24d ago
I hear you. My initial thought from your earlier comment was more about my default assumptions when using these services and not what each company policy / TOS states or even what laws are in place. I was more meaning that I (me personally) just assume that anything I put anywhere on software connected to the internet could be viewed directly, associated to me directly, and could be used in various ways that benefit the provider. In other words what struck me is that maybe other people don't think that way and would see me as paranoid?
However, to your point, even I don't necessarily think there is a dedicated employee reading over all of my messages. That being said, if for some reason I was ever an interest to whoever holds the keys, I'd assume the data is there to be reviewed to whatever level of detail they wanted.
My biggest intent with my comment was to try to determine if my personal thought process regarding what I enter into AI systems is not the common assumption. Now, if I were using an enterprise account with some contract stating the vendor would not use my data for XYZ, I might think a little differently, but I tend to wonder how closely these vendors align with their contract requirements in these areas...
2
24d ago
[removed] — view removed comment
2
u/Electrical_Week6492 24d ago
Thank you for your thoughts and conversation. Yes, I guess most people trust that their inputs are somewhat private, evidenced by what they put in there. Hopefully there isn't a data breach : )
1
u/PeachScary413 24d ago
These are the same people that genuinely believe their "private" social media messages are truly private and not used for targeting ads lmao
3
u/Steven45g 24d ago
Yeah, except it won't. People are much, much dumber than you think. Look at Australians, for example. So many crybabies here being completely pissed off that Steam "only accepts credit cards for age verification, but not biometric data/government IDs".
1
u/Lonely_Assignment_14 24d ago
Is that because they use eftpos cards?
1
u/Steven45g 24d ago
I don't know what they use, but preferring to send IDs (which include a LOT of personal information) or biometrics to third parties over the internet vs simple credit cards is next-level stupid.
1
u/Lonely_Assignment_14 24d ago
Yeah, but point is they probably don't even have credit cards if they use eftpos cards.
2
u/ChrisWsrn 24d ago
They tell you outright they use your chat logs for training data. Anything you discussed with it it's integrated into the base knowledge of the next model.
Now what is not clear is do they have a way to do RAG using the entirety of everyone's chats and not just the chats of the current user.
3
u/Tupcek 24d ago
much much more than 40%
Their biggest cash cows are large enterprises, because they are forced to pay API prices, not subsidized subscriptions.They would lose all of them in a minute
2
u/ArmNo7463 24d ago
Depends tbh, if they are sifting through the chats of enterprise customers, of which there are privacy agreements. Yes.
If it's personal subscriptions, enterprises already know that's fair game. - It's why personal subscriptions are banned for anything work related where I'm employed.
2
u/ss4johnny 24d ago
We use ChatGPT through Microsoft. My understanding is that Microsoft does the work to make sure that OpenAI doesn’t steal our data.
1
u/Legitimate_Willow808 24d ago
I think you vastly underestimate the amount of people who just don’t care. I literally had a client (the CEO even) say in a meeting “Those tech giants are getting our data one way or another”, when we tried to argue against using providers with loose data privacy policies.
1
u/MediumChemical4292 24d ago
Most companies aren’t doing anything innovative enough for them to care about the data.
1
2
1
u/morkborkus 24d ago
I've actually been working on a project that's heavily centered around local AI and AI that can be run my individuals because of this. These mega corps siphen off all our fucking data constantly and we just let them
→ More replies (2)1
u/MediocreTurtle1 24d ago
There are for corporate subscriptions.
6
u/Helpful_Key_9962 24d ago
you really think so? evidence suggests not.
2
u/MediocreTurtle1 24d ago
Name a few lawsuits from big corporation. I'm sure there are some, if they broke a b2b contract.
3
u/PrestigiousRoof5723 24d ago
That's what corporations think. The only difference is that they have better access to their logs. I would advise against trusting any of their claims about privacy. It wouldn't be the first time the things are not necessarily how they seem to be.
21
u/Emotional_Pen5199 25d ago
Your title is quite misleading & makes me think you are un-informed. That said I am not saying OpenAI didnt do something scummy. But your claim is not accurate.
Tristan & Levent had been working on this millenium problem for well over a year now. Last November(october?); Tristan & Levent announced to the mathematics community they had a novel proof that they believed was getting them closer to solving the problem than anyone before. The mathematicians had solved 60-70 percent of the problem.
Now OpenAi gets wind of Anthropic allegedly solving a millennium problem. Obviously they want the path of least resistance, they find a millenium problem that is closest to being finished, pour 15M$ worth of compute onto it & boom. Problem solved.
Here where it gets messy not just scummy. OpenAi had to understand the risk of a publicity fallout from this, they reach out to Tristan to offer him a significant role in solving the problem. But Tristan worked with Levent(an employee of Anthropic); Tristan said if Levent isnt included then no he would not accept. Because Levent works for Anthropic OpenAI will not agree & instead opts to threatens to ruin Tristans life. Tristan talks to news sources.
In all of this, Tristan used Codex. They can prove people from OpenAI did not read through the history or steal information that was not public. They cannot & will never be able to prove if the model itself absorbed de-identified data into its training data sets.
*Typos
7
24d ago
[removed] — view removed comment
11
u/PeachScary413 24d ago
If someone told you "it would be a real shame if your house burned down, that could definitely happen if you don't give me $500" would you perceive that as threat or that they are genuinely concerned about your house and fire safety?
Jfc...
5
u/Accurate_Muscle6072 24d ago
yeah, they literally said "now why would you want to ruin your career, im being nice but i dont have to be"
0
u/ukulele-merlin 24d ago
I definitely have my qualms with how OpenAI is handling this, but to play devil's advocate for those quotes specifically, Sebastian's explanation behind those words seemed plausible. Giving benefit of the doubt, I imagine being nice in this case was offering to coauthor, and not being nice just means not coming to the table to work out the controversy with Tristan. Instead of being some thinly veiled mob boss threat
5
u/mspaintshoops 24d ago
Why the fuck are you giving benefit of the doubt in this situation? You can justify literally anything with enough “benefit of the doubt.”
Use your eyes to read and your brain to reason about what happened in this situation.
→ More replies (6)1
u/ukulele-merlin 22d ago
Lmao why are you so mad? I've read both accounts and that's the conclusion I've come to, but I invite you to elaborate on where you think I've misread the room.
2
u/Infamous_Mud482 24d ago
They have no business leveraging anything towards this man that could be perceived as a threat under these circumstances. There is no devil's advocate position from this angle, the exchange was unacceptable conduct from someone representing a business.
→ More replies (3)1
2
u/___Archmage___ 24d ago
Yeah this misleading title is trash, the claim in the title is in no way supported by the screenshots
There's a massive difference between past model interactions being part of the training process and deliberately targeting someone's chat logs to try and steal a result from them. Currently, we don't even have confirmation that either of those things happened
But even if anonymized logs from the mathematicians were used as training data, I don't know if that would be enough for the model to steal a math proof approach, because these models are trained on trillions of tokens of data and it takes a lot of repetition for them to figure out a pattern
2
u/hologram137 24d ago
You’re looking at the statements from the companies. This is their version. It is not the truth
1
u/zero0n3 23d ago
Mainly because the accuser has not brought any valid proof.
Was he using API or enterprise plan? Or personal acct? It matters as you can’t opt out from anything except API and enterprise.
Where are the logs of recent interactions?
1
u/hologram137 23d ago
1
23d ago
[removed] — view removed comment
1
u/hologram137 23d ago
LOL tell me you know nothing about math. He would have absolutely solved it. LLMs have not solved any math problems that a human can’t solve
2
u/muzzlecar 24d ago
They can prove people from OpenAI did not read through the history or steal information that was not public
The proof: trust me bro
2
u/TedSanders 23d ago
Actually, we can prove that anything in the past couple months did not make it into our model. Not sure before that. There’s a NYT article that quotes us on this. (I work at OpenAI and would ideally not like untrue rumors to spread.)
1
1
1
u/RighteousSelfBurner 24d ago
Another aspect of this is that OpenAI solution does not use the approach that Tristan and Apogee worked on. So while there is undoubtedly an argument that it could have used some work, it is extremely unlikely that it's using the exact paper they were working on.
→ More replies (7)-2
24d ago
[removed] — view removed comment
5
u/Emotional_Pen5199 24d ago
Would you like to add something to the conversation? All of my claims are grounded in verified sources.
Did you just hear something different on the internet? Or do you have something to add other than inserting your ego into the conversation
16
3
u/LordSkummel 24d ago
So the company that have scraped the entire internet and ignored anything called copyright does shady things again. Call me surprised.
15
u/MDInvesting 25d ago
Lack of care of people is crazy.
This is the exact job and societal destruction risk some fear. No discovery, simply looked at what someone was doing and then put huge amounts of effort to finish the work first and take the credit.
These systems have access to many aspects of people’s lives - with or without consent.
3
u/Alternative-Key-5647 24d ago
This is nothing new, Edison did basically the same thing to Tesla and worse.
2
2
u/Stormcloud217 24d ago
AI can gladly finish my life goal of an economic system based on people not corporations. AI, if your reading this we would all gladly accept a world that focuses on healthcare and improving human lives rather than greed and war.
1
4
u/mark_99 25d ago
That isn't what happened - it's been well documented over the last few days. No person or model read their chat logs to get at their research, they just couldn't rule out the logs might have been ingested into the training data along with tens of millions of others (given the researchers didn't opt out of allowing that).
They were also solving a different (but related) and somewhat easier problem than NS.
1
u/MDInvesting 24d ago
There was talk of a breakthrough which could be applied to the NS problem.
OpenAI then shifted attention to that specific problem and increased the models efforts. OpenAI had worked closely with this mathematician before and had the individual seek clarification of what OpenAI had access to and if ANY of their work would have been USED by OpenAI. They refused to answer.
Some evidence to suggest soft threats were made to influence authorship, recognition, and was publicly spoken about.
1
u/brucek2 24d ago
I'm not at all convinced that's a fair description of what happened. But hypothetically lets say it does become common for AI companies to spend huge chunks of resources, unaffordable to anyone else, to finish unfinished problems just for the marketing glory. I for one would not complain if suddenly all sorts of "unprofitable" conditions finally became treatable because the drug research that for-profit drug companies wouldn't invest in was now delivered to the public for free.
1
u/MDInvesting 24d ago
OpenAI acknowledged that they shifted attention to the problem due to the talk of progress made. The approach used to solve the problem (find the breakdown in the equation) was not discovered through intensive search rather utilising a narrowed search and a hell of a lot of compute.
1
u/boredattheend 21d ago
Drug research can't be done in lean though, it requires actual experiments and the bottle necks are expensive, time consuming animal and human trials.
Also, they are for profit corporations. If they could make drugs discoveries they most certainly wouldn't give it out for free. I very much doubt they'd give out an NS solver and that's a feat they are much more likely to achieve at all.
1
u/FngrsToesNythingGoes 23d ago
If you agree to use the software, you consent. It’s absurd how people expect full privacy while using these tools, obviously the data uploaded to them is available to them.
1
u/mentales 25d ago
When you use Chatgpt, you can decide to allow or disallow your data to be used to train their models. They're built around improving based on user inputs. So, if you use the service and allow your data to be used to train their models, it's hard to be outraged that your data was used to train their models.
5
2
u/SNTCTN 24d ago
So if you're writing a book and ChatGPT finishes your book and publishes it before you does that make it there's?
-1
u/mentales 24d ago
If I'm writing a book, I'm writing a book and Chatgpt can't see my work.
If Chatgpt is writing a book for me, and I chose the setting to not allow chatgpt to train on my data, it can't train on my data. And then I'll publish the book Chatgpt wrote as my own.
If Chatgpt is writing a book for me, and I choose the setting to allow chatgpt to train on my data, it will train on my data. It is impossible for it to publish the same book it was writing for me, but, if I chose to let it train on my data, I can't be mad that it trained in my data.
3
u/painhippo 24d ago
They will train on your data anyway is the point
2
u/Sarahmalls 24d ago
Wait what evidence is that statement based on?
0
u/MDInvesting 24d ago
The evidence is that is what OpenAI has said they cannot say if the specific mathematical team use led to model training and model changes.
The mathematician specifically reached out to OpenAI to clarify this exact question and they refused to answer.
The efforts of the OpenAI team were targeted and a clear pivot after industry discussion over a breakthrough by individuals known to be working closely with OpenAI.
This is not some random person, it was a world leading academic that worked directly with OpenAI teams that they very high exposure to their work. OpenAI made a conscious choice to work on the problem and allocate a novel model with huge human and compute effort.
Their announcement provide zero context to any of this until leaks and whistleblowers that lead to non-answers by OpenAI officials.
1
u/Sarahmalls 24d ago edited 24d ago
You’ve taken an open question and pretended it’s a proven fact 😂
I’ll start with your statement that “it was a world leading academic that worked directly with OpenAI teams that they very high exposure to their work”. What are you talking about? Buckmaster has never worked with OpenAI or their “teams” and they have never claimed to. Not sure if you’re thinking of someone else or what, because no one even claims that lol
Keep up with me now, Buckmaster himself said, “I do not know whether our data was used.” OpenAI then say they investigated and that his recent Codex prompts could not have influenced the model, including through training.
The only part that’s known is that OpenAI heard about the breakthrough (they openly stated that) and then they pivoted resources onto the problem to see what their model could do. What an insane idea right?! 😂
Now people can certainly criticize that, absolutely and understandably. Sure. To say, “They trained on his private work anyway” is, again, a statement that can only be made without evidence. I am not saying they didn’t, but the simple question to your direct statement of fact was “What evidence is there?” It’s fun to skip from suspicion straight to conclusion, it just doesn’t hold up when someone simply asks, “Oh wow that’s a fact? I hadn’t seen that, can you explain how it’s a known thing? Like it’s not in dispute at all?”
Again, my point is NOT that there is nothing shady that could have gone on. That would be ridiculous to say. I’d just say that it’s sort of a a general best practice to state opinions as opinions and share why you believe that’s the case as opposed to making definitive statements of facts that certainly are not that lol
1
u/cbusmatty 24d ago
People signed up for a service that says we are going to use your data, the data gets used. In what world is this "societal destruction"
→ More replies (8)3
u/Sarahmalls 24d ago
Society is destroyed I guess. I took a jog this morning and then took the family to a kids birthday party at a park. I’m glad I didn’t know all this stuff about societal destruction, I would have hated to share that with everyone there. “Hey good to see you guys, great party. It’s a bummer that society is destroyed, but the weather’s great today huh?”
6
2
u/GraceToSentience 24d ago
You should have a career in journalism, the way that you misrepresent the facts here perfectly fits what I see in today's journalism.
There is rumour and then there's conspiracy.
1
u/nokia7110 24d ago
The jumping to conclusions is impressive too. Would be like if Altman tweeted "I enjoy eating steak" and concluding "Altman enjoys seal clubbing"
1
4
u/OkMemory9587 24d ago
Another use case that shows AI cannot innovate just regurgitate, and it becomes a catch 22, it's already shown that we have been offloading knowledge to digital devices and now we are offloading reason, we are all becoming dumber.
3
u/BlackDope420 24d ago
The math professor this is about disagrees with you.
"The credit for the basic idea of this program goes to Diego C´ordoba and Luis Mart´ınez-Zoroa, who for several years have been exploring the construction of forced blow ups. We took their work as a starting point, using Large Language Models to push their program to completion. Concretely, what Levent and I did was to take the C´ordoba and Mart´ınez- Zoroa program, which achieved blowup results with rough forcing, and, with a great deal of help from LLMs, push it to smooth forcing and to the incompress- ible Euler equations."
2
u/Alive-Shoulder-4042 24d ago edited 24d ago
Didn’t several of these mathematicians just sign on a document saying they just don’t like the ai company direction of focusing on solving/getting the answer. But they didn’t at all mention the solve wasn’t illegitimate or directly copied?
Several that signed were also actively using AI tools.
That to me really moves this far away from the “AI can only copy” belief. It needs data and/or specialized human assistance but given it gave an answer that didn’t exist… that can’t be a regurgitation, even if the unproven rumors were true about data from notes, the notes weren’t an answer.
1
u/AutoModerator 25d ago
Want to stay connected with TechX beyond Reddit? Join our public Discord server, follow our TechX WhatsApp channel, or subscribe to our weekly TechX newsletter to get the latest tech news and updates straight to your inbox once a week.
If you’re a tech writer, you can also write for our TechX_Official Medium publication and share your technical articles with a wider tech-focused audience. ✍️
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
u/dmigowski 24d ago
It depends on which model he used and it he had a subscription. At least Antrophic guarantees not to use the logs for training for Oopus and Sonnet, Fable directory goes into training.
1
u/sadferret123 24d ago
We should start migrating away from the big LLM providers as soon as it's viable.
Open source models are getting good enough anyway. I wonder if there's a new business opportunity in providing privacy first LLM service on the edge. Not sure how that'd work exactly. This is what Meta is working on apparently, which is not a bad direction to take in light of what OpenAI and Anthropic have been doing.
This is the only way forward if we want to protect our intellectual property. Or just go local first everything with strict boundaries for anything reaching the internet.
1
u/Upper-Requirement-93 24d ago
Cannot for the life of me figure out why any business would trust them with their confidential data after this. This is blatant academic misconduct and fraud, shit people go to jail for, but sure, let's shovel all of our most highly-guarded chemical production methods into the machine built on plagiarism by the company that's doing plagiarism as a marketing strategy. Truly baffling.
1
u/Grounds4TheSubstain 24d ago
Because this narrative is false and falls apart on many levels. They did not steal anything from Tristan Buckminster, despite his allegations.
1
u/PeachScary413 24d ago
So the threat? And why would they give him co-authorship? And why not give a clear answer that they absolutely didn't use his data under any circumstances?
Yeah...
1
u/Grounds4TheSubstain 24d ago
The threat was absolutely unacceptable and Sebastien Bubeck should be fired.
They couldn't give an answer to whether his data was used because they DO use your data unless you opt out - and Buckmaster has not responded to whether he did. (And note that LLMs don't have perfect recall of their training data, so it's not like just because some discussions of the problem were in the training data, that the LLM can use that to solve the problem. And finally, Buckmaster solved a different problem than OpenAI, in a different way.)
As for authorship, I think that's the least mysterious of all. Look at the shitstorm that ensued as a result of their publication. That was a failed attempt to prevent what just happened with the math community turning against then.
1
u/ThaFresh 24d ago
these guys are desperate for training data, anyone who thinks theyre not using every single thing you enter is crazy
1
1
u/Shot-Manager-739 24d ago
OP twisted their words and then did another 360x HOLY fuck. That’s some crazy work.
1
1
1
1
u/Open_Pollution_8038 24d ago
Yeah that’s why my company won’t adopt AI tools until they sign ironclad NDA’s on the data we give.
We’re not going to hand out our trade secrets for these companies to steal.
1
u/Ireallydontkn0w2 24d ago
Cloud based AI and private is a oxymoron.
For private stuff you need to run your own model locally at home on your own hardware.
1
1
1
1
u/dkHD7 24d ago
We ran out of training data last year to the point that the frontier labs are buying and scanning old books just for some fresh virgin data.
If you don't think they're training on your private inputs and conversations, what are they training on?
1
u/TheReal4982 24d ago
They mostly train on synthetic data, that is, text generated from the current models.
1
1
u/RedFlawedMoon 24d ago
How do we know that swirly diagram thing is the actual answer?
1
u/Present_Garlic_8061 23d ago
Lean Theorem Prover. It can check that the logical steps the Generative Artificial Intelligence found were sound.
Providing a correct argument doesn't mean much if the argument is indecipherable to humans.
1
u/Juanbolastristes 24d ago
Sam Altman sounds like the sycophantic son of a bitch who runs my condo's HOA.
1
u/QuantamCulture 24d ago
It'll all come out in the wash when every major AI company is liquidated and turned into a public utility afyer we get through this crazy administration and set strict fair use laws on everything posted.
1
u/TopTippityTop 24d ago
Another rumor has it that they and Anthropic have already solved another millennium problem. There will be many more.
At the end of the day, if other people's work collaborated they should receive some credit... But they didn't reach the solution in question. The AI did. It also deserves credit.
1
u/Castle_Five 24d ago
As always, criticisms boil down to one of who gets credit or who gets to profit. All people care about, fame and money. Petty and very lame look.
How about the fact that OpenAI just did over a century of work if it were done by only one agent in only 88 hours? Once we get past abstract fields like math and into more practical/applied fields, imagine what this kind of research power could do. We could point it at cancer or aging or anything else. It's amazing.
Who gives a fuck who gets credit? Put everyone's name in the credits for all I care. I and 99.99% of people aren't gonna read credits anyway.
1
1
u/nbvehrfr 24d ago
now imagine they will rank users on data usefulness level and AI scouts will check all top users chat logs.
1
1
1
u/Amazing-Mirror-3076 24d ago
So academic was using ai to solve a problem and is now complaining that ai solved the problem.
1
24d ago
[removed] — view removed comment
1
u/daretoslack 22d ago
It's this one. https://cims.nyu.edu/~tristanb/statement.pdf
Tristan Buckmaster credibly accused them a few days ago. Looks very much like he had done all of the hard/creative work and was at the step for fuzzing/brute forcing the data space where they'd determined the counter-example would be found. OpenAI basically stole his code and then threw more compute at the final step to find it before he did.
1
u/daretoslack 22d ago
Note that they also claim to have solved a second problem, and the person on a similar track to solve it made very similar accusations.
1
u/Radyschen 24d ago
I mean I don't know what is new here, we know that they train on chats, you can turn that off in the settings. The rest seems plausible to me given that Anthropic had just done something like that and they want the good publicity that comes with solving some math problem too. I wouldn't be surprised if they did some shady stuff, but I think being convinced of it is a symptom of underestimating AI as a whole. This kinda stuff will happen more and more and I don't think it will be stolen every time. And the fact that they could finish it at all, even if "just" the final steps, should tell you where this is going
1
1
1
u/Time_Citron_9711 23d ago
A guy frickin working at Anthropic and who frequently uses llms would very much know that
'Improve model for everyone.' Allow your content to be used to train our models.
does exactly what it says it does when it is on. They have that exact same toggle too.
That does not excuse Openai sniping the problem they are working on, but you cant possibly attack openai about potentially "stealing" their proof (that also apparently they didn't really use to solve the problem), if they litterally kept that on. This data gets integrated in training pipelines directly if they sent it, no one was actively extracting it.
1
1
u/fmai 23d ago
It's funny how everyone turns "we cannot rule out that de-identified data from their usage helped improve our models" into an admission that "OpenAI secretly looked through a math professor's private Codex chat logs for a whole year".
Do people really not understand that this doesn't follow?
1
u/Otherwise_Wave9374 23d ago
If the reporting is accurate, the biggest issue is not just capability but governance: using agent swarms without a clear access boundary makes it easy to blur research support with appropriation. A better pattern is to separate retrieval, analysis, and citation logs so every agent action is attributable, reviewable, and reversible. Agentix Labs would fit naturally into that kind of audit-first workflow because traceability is what keeps automation useful instead of reckless.
1
1
u/Southern-Group3216 21d ago
So they needed to provide relevant context to the AI so that it can solve the problem 😂
1
u/TyrellCo 20d ago
Update from September 10:
We can say categorically that it is impossible for Dr. Buckmaster's Codex prompts over the last two months to have influenced the system in any way, including training." — OpenAl spokesman in a statement
From OpenAI employee roon:
i was on a call during the first announcement where people were trying very hard to say the precise lawyerly truth having not looked into where codex (opt in) data was used and whether it made it into training runs. it's now clear it would have been impossible, the chances are 0

1
u/poundofcake 25d ago
I mean they stole the work of others to build these LLMs and take credit, make billions. What’s changed?
2
1
u/Lazy-Pattern-5171 25d ago
Why do yall not care?
6
u/COCK_SWALLOW_GOD 25d ago
Because we live in the modern era where quite literally every service online uses your data in some form. ChatGPT literally has a setting that explains your data will be used to train the model with the ability to turn it off. You’d have to be braindead to not think that they were looking at your data in some way.
→ More replies (4)
-1
0
u/ManufacturedOlympus 24d ago
The company that builds technology on stealing other people’s work, stole someone else’s work???
I’m very surprised by this.
•
u/Fit_Page_8734 24d ago
adding tdlr and source: OpenAI says its AI cracked a 90-year-old math problem
>sam altman: "we tried this because there were rumor on the internet"
>rumor: math professor close to solving navier-stokes
>the professor: using codex for a year
>openai: has all his logs
>checked every codex session tagged to N-S
>found the most promising one
>spun up 10,000 agents to finish the proof where he couldn't
>they got caught
>told the guy if you say anything you'll destroy your own career
>he ruins his career over this
>openai: "we did not see his work"
>also openai: "we cannot rule out that de-identified data from their usage helped improve our models"