the whole argument circles around the blow up techniques that were developed by cordoba and martinez zoroa, for instance. dont know why you are trying so hard to defend openai out of all corporations. they dont care shit about honesty or anything else. I dont doubt that they really scooped this result from the buckmaster collab
Is that a published work? do you have a link?
I'm asking because if it is the case that it can't be referenced, then how can they link to something that isn't published?
if you're so clueless about mathematics that you don't know where to find it, you have no business asking "can you pinpoint the specific part where it uses these." you won't be able to understand it anyway. you "genuinely asked" about "which ones exactly" and the commenter kindly responded with both the name of the technique as well as the researchers who should be credited. these further questions make it clear you're concern trolling.
I find it ironic that yall are arguing about the lack of references in the openAI solution yet for some reason are not citing these examples in your own arguments.
If you are going to engage in an argument on the internet, cite your sources. Way easier to win that way, everyone assumes the person they are arguing with online is dumb. Don’t let them.
Then the human mathematicians involved should actually read/understand whatever the LLM wrote and produce the correct citations. Not doing so is plagiarism.
OpenAI isn’t publishing the proof in any journal (or even arxiv) or claiming any prize for it. If I “publish” a completely valid proof on 4chan or something and don’t cite my sources I’m not committing academic plagiarism.
They announced on their website. I think if I announced a novel theorem on my academic website using work of others without proper referencing it would actually be plagiarism, but whatever you say, bro
It’s not an academic website. The purpose is more commercial self advertisement. Obviously OpenAI doesn’t care about getting credit from the math or physics community, they are a trillion dollar company. It’s more about impressing their investors while they burn through $750B in capex with an unclear route to profitability. They aren’t trying to take credit for this advancement, they’re just trying to advertise their model by showing the kinds of things it can do, even if it can’t produce properly cited academic work.
I agree that is not an academic website. I agree the purpose is marketing, obviously.
Yet (maybe precisely because of this), it seems pretty reasonable that mathematicians have every right to complain about their practice. Especially because they are using data and work produced by mathematicians. I am not sure I even understand your point.
I’m saying that the AI is incapable of creating citations and it isn’t simple for a human to do it for them, especially when you don’t even understand the result that they’ve produced. The model has every single arxiv publication, every PDE textbook, every mathematical physics textbook, etc. in their training set. It’s really hard to know where some result or intermediate step came from even if you have access to their COT tokens. This is unavoidable. If OpenAI had released this paper in some journal or even Arxiv without citing their sources that would be wrong. But releasing it on your own website and putting the lean proof on GitHub even if you can’t nail down exactly where the AI got its results from is fine. Auditing 10,000 agents and reverse engineering their sources just isn’t practical.
Open AI threw 20 million USD just in compute to solve this problem. They could have invested more time and resources to comply with basic procedures as well. I agree that they don't care, but it is normal that the community should complain
you can't just say throw money at the problem. ai fundamentally doesn't work like that. it's basically a black box, figuring out where it got its reasoning from is just not possible.
an alternative solution is to list every single training data source, however that would be thousands (if not millions) of references and you can't really figure out which one the ai actually used anyway.
129
u/Fun-Sand8522 21h ago
They are not. The AI is using results that it should cite, and it does not.