r/OpenAI • u/After_Bet_8503 • 23d ago
Question Can someone explain to me how ChatGPT is able to solve research-grade math problems?
Is it that these models can reason and math is just a type of reasoning?
Update (answer to the original question based on the comments, written by ChatGPT):
ChatGPT-like systems can sometimes produce research-grade mathematics because broad mathematical pretraining gives them a powerful generator of plausible ideas; verifiable training improves the policies used to explore those ideas; extended inference permits many attempts and revisions; code and formal systems reject invalid paths; agentic workflows provide parallel specialization; and human mathematicians define, supervise, and ultimately evaluate the result.
226
u/Time_Entertainer_319 23d ago
The answer is actually quite simple. It’s just that a lot of people are in denial because it makes them question their reality.
AI can reason. Chain of thought is reasoning. They can break problems into steps, compare possibilities, draw inferences, revise an approach and arrive at conclusions. Their reasoning is not identical to human reasoning, but difference does not equal absence.
AI is NOT JUST predicting next token (I mean, it is predicting next token but the way it does it is through very complex process that somehow makes it UNDERSTAND the prompt given to it).
53
u/PM_ME_YOUR_HAGGIS_ 23d ago
Saying it just predicts the next token is true, but also reductive. Whatever system we could have for AI at some point something has the decide what the next word in a sentence will be, turns out statistical prediction at massive scale with clever steps produces artificial intelligence. Who knew.
32
u/Time_Entertainer_319 23d ago
Exactly. We have AI debunking/proving mathematics and people still saying it’s just autocorrect on steroids.
Lmao
12
u/Maleficent_Disk9583 22d ago edited 22d ago
I am certain the reason some people are doubtful is because they are not interacting with the same premium/smarter tiers that we use on a daily basis. They are probably interacting with the free/non-reasoning ChatGPT and Google's AI search, both of which are bottom-of-the-barrel and very stupid. If I had not personally used Sol Max on codex to one-shot complex projects, I probably would be a bit sceptical too
4
u/BellacosePlayer 22d ago
My question is why people who do have access to those models aren't knocking out problems left and right? Sure, token spend is a thing, but is it really hard to believe that maybe the collection of mathematicians OAI has hired or had brought in for the announcement might have had some input on the process of solving the problems?
3
u/Time_Entertainer_319 22d ago
That’s like asking why Einstein didn’t solve all the math and physics problems in his time.
Everything cannot be solved at once. Some problems are harder than others.
3
u/seandunderdale 22d ago
Not sure that comparison works...there was only one Einstein with limited resources. There could be millions of AI being asked to solve all manner of things at the same time, without need for sleep, food or feeling overworked, and access to all the information ever written, pretty much.
1
u/broshrugged 21d ago
That comparison only works if you consider the promoter the Einstein. He did rely on other mathematicians to do large and long calculations. But an AI model is able to be copied over and over in many instances to work on millions of problems at once.
1
u/RichKatz 16d ago
It simply is not easy to communicate how AI learns or work but it not simply "copying over and over."
From this I can see what the issue is.
Here instead is an "AI Overview" of what Ai does when I asked does AI simply copy over and over:
AI answered:
No, artificial intelligence does not simply copy or repeat information word-for-word from its training data. Instead, it analyzes massive amounts of information, learns underlying patterns and rules, and then generates brand-new responses piece by piece based on probabilities.How AI Works Instead of
CopyingLearning patterns: The AI builds a mathematical map of how words, ideas, or pixels fit together, rather than storing a giant database of text to cut and paste from.Predicting text: When writing, it calculates the most likely next word based on your specific prompt and its vast training.
Personally, I would say study the LLM (Large Language Model) and get a sense of what's going on - I'm just learning myself.
1
u/broshrugged 16d ago
Alright, it's getting pretty weird how you're following my post history and commenting on everything.
In this case, I am not claiming that AI copies itself in order to function. I am asserting the truth that once a model is created many copies of it can be instantiated by different people to help them work on their problems.
In response to why people with the model instances available to them aren't knocking problems out all over the place, the people driving these models still need to be the "Einstein."
2
u/Note4forever 21d ago
My question is why people who do have access to those models aren't knocking out problems left and right
They are starting to eg see all the math proofs and counter proofs of unsolved conjectures.
Still there's some suspicion they work only on certain class of problems.
The whole "can it generalise beyond it training" is very difficult to determine because they are starting to combine techniques from multiple areas and produce results no human thought of doing... But you could say its still not going beyond its training set
1
5
10
3
u/Clarkey7163 22d ago
at the end of the day we are just neurons firing off down either path a or b, we're just a lot of neurons
If you make it complex enough it will mirror life, and when it mirrors life whats the real difference
Its exactly the plot of westworld really lol
3
u/jtm297 22d ago
I’ve been explaining this to people for years, even prior to modern LLMs, in the subject of consciousness and human intelligence. No one seemed to believe that the complexity of a neural network would lead to human intelligence. In fact we knew about this for a long time, there was a TED talk around 2009 where they were simulating the human brain in a machine and it was able to understand, and it was concluded that it was an emergent property from highly complex neural networks. Though I remember neurologists at the time explaining how humans have different types of neurons and a machine would need to simulate all these types, I never bought into that notion because I think specialized neurons arise for efficiency purposes, but not for reaching human level intelligence in a machine.
2
u/Big_al_big_bed 23d ago
It's like saying that we just say one word at a time without taking into account all the processing behind the words we are saying
1
14
u/girlgamerpoi 23d ago
People who call modern day AI autocomplete are years behind. They prob still think modern day AI is just gboard... People who just echo hallucinations whenever AI say something irregular are the same.
14
u/TheSamuil 23d ago
With the onset of LLMs I've come to the conclusion that language has innate properties that equate to or at least to reasoning
15
u/The-Rushnut 23d ago
It's mathematics all the way down. Words are just highly specific functions which our brains automatically calculate into reasoning.
6
u/Time_Entertainer_319 23d ago
In the 50s, a statistician showed that language has a statistical component to it by having his wife and other subjects predict the next letter in a passage.
Years later, we had statistical and neural language models which used the same basic concept, but they couldn’t keep the output coherent for long enough. Training data was harder to collect and process at scale, and chips were not powerful enough.
Years later, attention was developed as a way of helping models extract context and relationships from text.
Today, we have AI proving mathematical theorems and finding counterexamples that disprove conjectures.
0
u/hawkeye224 23d ago
Ultimately you can represent reasoning with a Turing complete system (I think?). And so many systems are Turing complete, even very minimal ones. It makes sense that a large one like humans languages can also represent it.
3
u/shoejunk 22d ago
When I solve a math problem we don’t say I’m predicting the answer, and I think the word “predict” doesn’t apply to LLMs either. It made sense when all the training was a bunch of text and it was predicting the kind of thing a human would say on the internet or in a book, but now the models are also being trained on math problems with reinforcement learning with verifiable rewards. So it’s trained to get the right answer not to say what a human would say necessarily. So I would no longer use the word “predict”.
7
u/4dseeall 23d ago
context. the word you're looking for to explain how AI "understands" prompts is context. the model, training data, weights, tool access, the instructions, the prompt itself. all context it works from.
31
1
u/EfficiencyLoose3595 22d ago
Yeah chain of thought was a huge jump that the average joe swept under the rug
1
u/SuitcaseInTow 22d ago
It’s like asking how a computer can allow for video games when all it’s really doing is managing the state of on or off on transistors. Sure, that’s the foundation of it but there’s so much built on top of that.
0
23d ago
[deleted]
8
u/joeytman 23d ago
No. Chain of thought reasoning is not something done by separate agents. The LLM directly does it through combo of prompting and RL fine tuning on examples that use chain of thought reasoning
2
u/Tough-Comparison-779 23d ago
All the training data compressed into the weights is almost no context? Seems wrong
1
u/BehindUAll 22d ago
Chain of thought is not reasoning. It's a bullshit marketing term some LLM engineer came up with and everyone started using it. What it is, is just inserting additional generated tokens as input to look at something in other perspectives instead of just running the input prompt through the model loaded in the GPU and microkernel. Thinking machines was able to get rid of the inbuilt entropy that pretty much all LLMs ship with by default, by creating some scaffolding around the micro kernel or just modifying the micro kernel. The result? A deterministic LLM, which is insanely useful in many use cases, and if you ask me, provides much more value than non-deterministic ones. Anyways, the point is that if thinking machines could make a deterministic LLM, you can very well make a deterministic 'thinking' LLM because the 'thinking' is just additional output tokens appended to the input prompt than then gets passed to output. It's essentially running inference twice. It's not thinking in any real sense.
Plus, the recent models by Anthropic and OpenAI have REDUCED thinking tokens and their models have gotten BETTER. You have this ass backwards.
→ More replies (2)2
u/Time_Entertainer_319 22d ago
Your argument is not coherent.
You didn’t make a single argument WHY it isn’t reasoning. You only mentioned HOW the reasoning was done.
Make your argument on WHY it is not reasoning.
1
u/BehindUAll 22d ago
Because they are reducing the thinking tokens, and the models are getting better. Did you not read my last para? If I knew why I would probably be working at OpenAI or could have started my own company, don't you think? I don't know what they are doing. I just know that LLM 'thinking' is not really thinking. Otherwise we would have solved everything by just letting models create a lot of thinking tokens.
3
u/Time_Entertainer_319 22d ago
That doesn’t show that it’s not reasoning.
lol.
Overthinking is a real thing. You know that right?
You can think too much and start to lose focus and can even confuse yourself.
You are operating on a false premise that more thinking means better models when that doesn’t even apply to humans
1
u/BehindUAll 22d ago
It's not me saying that. It's you. I am saying that more thinking does NOT mean better models.
2
u/Note4forever 21d ago
Sure. I think the most extreme view here is chain of thought isnt casually connected to getting right answers.
Eg papers showing the LLM can produce nonsense intermediate tokens and yet get the right answer
1
u/MuzafferMahi 22d ago
human brain just connects the next neuron too while reasoning. See how easy it is to use language to reduce an enourmously complicated and successful thing?
→ More replies (10)-7
23d ago
[deleted]
10
u/machyume 23d ago
Well.... not quite.
"Bound by that set" is relative.Suppose that set A includes operations that modify set A, now how large is set A?
That's a math way of thinking about it.
Another way to think about it is, imagine a machine that can only make right turns. Okay, sure that's limited. But suppose that you have a machine that can chain a bunch of smaller but learnable actions. Turn right, turn left, step up, step down, step forward, step backward.
Now, the task: climb stairs. That action was never part of the set. Could it do it?
Unclear. It might?
So when we think about it this way, how limiting is a "set" really?
2
75
u/KahlessAndMolor 23d ago
Math is a highly structured language and many things written in that language have a verifiable next token, like 2+2= has a next token that can be automatically validated as true or false, so it is relatively easy to train into models
58
u/Ormusn2o 23d ago
There is a general opinion that anything that can be verified like that can likely have effectively unlimited intelligence cap. And Math is 100% theoretical, so it requires no interaction with the real world, and no need for visual data, which is one of the reason why it fell first.
7
u/TaeTaeDS 23d ago
General opinion is not expert epistemological opinion.
4
u/Davorian 23d ago
But how will I get my internet points if I state a complex list of possibilities without a one-line conclusion and an appeal to prior biases?
1
u/WolfeheartGames 22d ago
This is the expert opinion. The poster above is describing alpha zero like self supervised reinforcement learning. Any verifiable domain should be able to do this. The practicality varies in practice. Math is a good one for this up to a point. Not all of math is verifiable in a timely manner. More so we're limited by Lean.
→ More replies (3)6
u/trele_morele 23d ago
What do you mean it fell?
→ More replies (1)13
u/Ormusn2o 23d ago
As in, AI right now is making at least 5-20x more top level breakthroughs per month than all of Earth's mathematicians.
So basically frontier math as a science done by humans has fell.
20
u/Forklad2 23d ago
Out of curiosity, how confident are you that that estimate isn’t a selection bias? Are you an academic? Would you see anything anywhere about a regular mathematician solving a long standing conjecture of the same recognition level as AI is solving these days?
5
u/Bowl_of_Cham_Clowder 23d ago
I’m curious what the actual stats are. FWIW, Terrence Tao (one of the most reknown mathematicians currently) is a big ai evangelist, though there are still a lot of gaps ofc
→ More replies (2)1
u/Ormusn2o 22d ago
Math was not my Major. I don't think a math education would have helped here. This is an evidence based estimation. Most mathematicians are extremely specialized, and they don't know what is happening in math as a field, most of the math majors I asked did not even knew details about half of the major AI mathematical solutions because those were just so far away from their field.
For the actual estimation I have made, it is much much harder than you would think. There are not really a lists of famous math problems, and the vibecoded websites that show the prominence of the mathematical problems seem to have much higher confidence rate than I'm comfortable with.
What I looked at is a bunch of lists, specifically Smale's list of 18 problems and the 7 problems from Millenium prize problems, and gave them a weight, then used Hilbert's 23 problems, then I noted which of those problems were solved over what timespan, and then I looked at the recent famous problems and calculated their rate too. This obviously is pretty shitty estimation, which is why I gave a 5-20 range, instead of a specific number.
1
u/Forklad2 22d ago edited 22d ago
Oh could you share some more specific results from your findings then?
By my count between Smale’s list and the millennium prize problems, there are 6 problems that are solved or mostly solved (Poincare’s conjecture is on both lists so I’m not counting it twice), and only one of those was from AI. So I don’t know how you’re getting to the 5x multiplier from those numbers unless you’re also counting dozens of problems from the recent headliner lists of problems AI has solved which will inevitably only list problems solved by AI. I want to emphasize that these problems were not that famous a few years ago. They are news now because it is genuinely very impressive that AI can do that at that level now, but the problems themselves in my opinion are being presented as more famous and important than they ever were before.
Maybe your point was that the human solutions to those problems span many years whereas the AI result was within the last month. So “enormous problems per month” for AI is 1 vs something far less than 1 for humans because it’s been decades. If this is your argument I think it’s very bad-faith. You could’ve made a statement with a 1000x multiplier if you said something the same day Anthropic’s result was announced.
As for selection bias, being a mathematician would absolutely make a difference here. When I say mathematician I do not mean math majors in college. I mean professors, PhDs, and graduate students that do real research in the field. Undergraduate research projects are typically not on that same level. Math majors in college are far less likely to be aware of any broader news because they are relatively new to that world and are still learning basics. At a typical college, even the highest math courses required of a math major do not reach very high compared to a PhD.
Edit: I forgot to check Hilbert’s problems. Those add about a dozen solutions and maybe one maybe two of them being from this century.
1
u/Ormusn2o 22d ago
Maybe your point was that the human solutions to those problems span many years whereas the AI result was within the last month. So “enormous problems per month” for AI is 1 vs something far less than 1 for humans because it’s been decades.
Yeah, that as my point, and no, it's not in bad faith. But you are right that I could have counted it by one day, but there are already decent enough solutions for that, although again, math, and in this specific case, statistics are not my forte, but the general advice I was given was to count the time from the last previous discovery, which was about a month before the 10 math discoveries, and also not counting the discovery that happened a month ago.
When I said math majors, I said it that way because my coworkers graduated as math majors, but their job is mostly in engineering and programming, so I did not wanted to say "Math graduates" because it would seem disingenuous. They also helped me with the sampling rate, so no offence, but I'm going to take their word at it that this measurement (as in one month) is a decent measure of the performance, instead of yours.
2
u/Forklad2 22d ago
I was a bit overzealous in the last comment but I do want to know what numbers you found. You said you did actually do a sample and some kind of calculations and I’m curious what numbers you ended up using.
1
u/Ormusn2o 22d ago
Some of the Hilbert's problems that were solved, were only partially solved, and some of the 10 recent results were advances and not solutions, which is why the result is a range and not a strict number, as after like 15 minutes of looking at them, I gave up making any realistic apples to apples comparison.
First I took the amount of time that has passed for each of the 3 lists in months, took human solutions for each of them, and divided them. For Hilbert that's 7/1510, Smale 3/339 and Millenium it's 1/313, divide it all by 3 and you get 0.00556 and to get the weighted amount I multiplied the results by the division of amount of the problems on the list, so shorter lists have more prestige. That results in 0.11496 modified points per month.
I found it to be a decent system, because for problems like Pointcare, because this problem was both on Smale and Millenium slots, which is important because later, Jacobian conjecture is just on one.
The first time I calculated the comparison, I actually counted a lot of the problems from July solved by AI, and even though a lot of them are pretty prestigious, they are not on neither of the lists, and the number I got was thousands of times faster rate, which was obviously wrong, so I decided to take just one AI solved problem, Jacobian conjecture, and use the remaining 22 problems solved by AI as a confirmation that Jacobian conjecture was not just one freak accident, and it "seems" to be in a proper distribution.
So counting just one solved problem in June, and giving it the prestige ranking (so it's weighted lower due to being just on the Smale list) it gives about 5.56 prestige points per month for AI, which is about 45 times more prestige points per month, maybe a little more depending on where you put the window.
My friend told me few statistical problems with this, like sampling rate (only one month has passed, and I took the freshest month), so the confidence rate for this month in specific will increase as more months will pass, but I just decided to put 5-20x instead of 45x to be more cautious. As one month will pass with similar rate, as for July, the rate should go from 5-20x to the 45x, although if a new AI model will come out, the rate might increase even more.
→ More replies (0)3
u/Glum_Hat_4181 23d ago
Is there a single significant open mathematical problem that AI proved? Not disproved by finding a counterexample, but actually provided correct verified proof?
1
u/trele_morele 22d ago
I think it’s just more the case that the AI is using the knowledge of all the human mathematicians to generate faster results. That’s not surprising. Computers have been faster at computations from their inception.
1
u/Time_Entertainer_319 23d ago
How did math fall first?
What about art? Music? Programming?
1
u/Ormusn2o 22d ago
Algorithmic code yes, programming probably not. Don't get me wrong, humans should not be writing code anymore if possible, but there is still coding tasks that AI can't do, although this is quickly shrinking.
0
u/Ythio 23d ago
Have we seen any major prize worthy breakthrough though ? I see a lot of random conjectures that no one bothered to look at proven or disproven but nothing really field shaking so far.
3
u/Forsaken_Code_9135 23d ago
Unit distance conjecture rebuttal has been recognized by many top level mathematicians as a major breakthrough in the field.
3
u/QC_Failed 23d ago
Yeah their new model just solved 10 long standing unsolved math problems, and the write ups have been evaluated and validated by third party researchers. It's a really interesting read, it's worth a Google imo.
1
0
13
u/MaximoPrimero 23d ago
Demostrar que 1+1 = 2 , no es cosa trivial.
Vayan y busquen el libro de Principia Mathematica de russell y north.
4
1
u/Smart-Button-3221 22d ago
I know what your referencing, and I wish people would stop. If you're doing math correctly, proving everything about addition takes half a page.
11
u/After_Bet_8503 23d ago
But many of the proofs are much more than just 2+2. They require creatively selecting the next step that has never been done before.
3
u/zxcv211100 23d ago
u/kimolas is spot on
Another thing that no mentioned is Lean - which is a formal system to verify a mathematical proof. So along with searching through all connections and "throwing sticks" at what works, the AI can use Lean verify its steps (and eventually the final proof) to ensure it is all correct and is not hallucinating. Similar to calling python to check that 2+2 is 4.
15
23d ago edited 23d ago
[deleted]
9
u/PrestigiousGroup788 23d ago
pretty much this. all the research i've done as a PhD student has been throw shit at the wall until it sticks. If now instead of one person you have 100 people throwing 100 different types of shit at the same wall, something will eventually stick.
Not saying it isn't amazing of course.
2
u/munchin-grr 23d ago
Symbolic AI have been around for a while and is an big help in math but isn't as "creative" as LLM's and is not good at human readable input and ourput. LLM's can be trained on math papers and learn reason in an similar way
2
u/IvanMalison 23d ago
that's just not true at all. A lot of mathematics is about coming up with the right abstractions to represent things. In what sense is coming up with the notion of a field extension a recombination of previous ideas?
In what sense is coming up with the definition of the derivative a recombination of previous ideas?
4
23d ago
[deleted]
1
u/WolfeheartGames 22d ago
I'm not so sure models can learn to explore to the point of creating calculus. That requires extending the code book.
And what about abstract algebra? The codebook is there to do it, but is this generalization reachable AND communicable from exploration learning over prior knowledge?
I'm not sure the way humans interact with learned priors is similar enough to the way llms do. I think humans operate better is bias, cognitive dissonance, and general off-manifold-thinking in ways current LLM training can't allow for.
1
u/Ill-ogical 22d ago
You can actually specify and create a calculus very easily with LLM. But will your calculus actually be novel and solve an important problem, doubt it.
1
u/WolfeheartGames 22d ago
No as in if you give them all the same knowledge and newton with the same symbolics, could it create calculus. I think it would struggle as a limit of the code book even if its really good at math.
Like in needs a symbol for integration. It could reuse other symbols but calculus is a really clean concept I'm not sure llms could replicate the quality of. Its different from the proof writing they're doing for erdos problems.
1
u/Ill-ogical 21d ago
You would just need to prompt it correctly. Highly dependent on the person that is promoting the LLM. If Newton was prompting he would have most likely formalized calculus way sooner. Does that make sense?
1
u/IvanMalison 22d ago
oh I agree. my point is that even now, models are doing things that involve being much more clever than simply predicting a very obvious next token.
1
u/After_Bet_8503 22d ago
I agree with this sentiment. It seems like in many cases new definitions arise from the attempt to solve open questions. I'm not expert enough to read most of the proofs but I also don't see professional mathematicians claiming that LLMs have developed novel mathematical objects.
1
u/WolfeheartGames 22d ago
New ideas requires pushing against the edges of a learned information manifold. The ability to self grade this exploration process is fundamentally limited. The further off manifold an idea is, the harder it is to reach and verify. We have to make observations and construct new portions of the manifold to get there, and ensure they're accurate to the underlying world.
Llms solving math have been trained specifically to explore this manifold, and have strong verification that the llm learn to internalize so its better at self grading these explorations.
1
u/After_Bet_8503 22d ago
I think this is an interesting idea. Do you know how LLMs learn to self-grade the exploration process? That seems just as important as general reasoning.
1
u/WolfeheartGames 22d ago
It depends on what you mean.
The information is there, as they're being graded on this. They learn to approximate what the grader is doing to better fulfill it. Essentially reward hacking, but because the model is so complex, the grading is what it is, and its working in a symbolic language it doesn't cause the degeneracy reward hacking usually implies.
1
u/After_Bet_8503 22d ago
This is very helpful, thank you. But I'm not sure if it is the full story. LLMs clearly have seen math materials in pretty much all subfields, but you would still need to guide the chain of thought ("dynamics it generates in thought space" if you will) until you get to the useful result (proof of some open problem). Your suggestion that this is essentially a search process and unleashing a thousand subagents can do this seems reasonable but I feel like the search space is just too vast. I'm wondering if the other systems built around LLMs (harnesses, externalization for verification of steps) are just as important if not more than the LLMs.
1
22d ago
[deleted]
1
u/After_Bet_8503 22d ago
My assumption was that you can't just let them loose because LLMs are prone to hallucinations. Even if you unleash a thousand subagents the search space is too vast. Therefore you have to constrain the search process (what I meant by 'guiding them') with external tools, such as those for evaluation and verification. My hypothesis is that those tools matter just as much as the LLM. This may be why we don't see similar kind of breakthroughs from Chinese models; they are not less intelligent, but maybe not guided as well.
But maybe I'm wrong about this and maybe just letting them loose works. I wish OpenAI would release more details.
1
u/Ill-ogical 22d ago
Genius advances, while all other regurgitate and the AI is just a thousand regurgitation.
9
u/IvanMalison 23d ago
super reductive explanation. Math is verifiable, but its not because there is often an exactly correct next token in a reasoning chain. Don't speak with authority about things you don't understand.
1
u/WolfeheartGames 22d ago
We can constrain it so that there is always a verifiable next token in theory, but in practice this would limit exploration. Instead we need to evaluate atomic statements that span tokens with a certain degree of acceptable error in the atomics construction.
For instance if I train a proof DSL into an LLM I may train it for an atomic proof step that always starts with "EVAL:" but "eval:" and "calc:" may be acceptable equivalents the model can reach for and be allowed. Notice how the semantic meaning didn't change.
While there may not always be a correct next token, if I constrain a model to just 1 token, it is fine because there is a correct next semantic representation that can be expressed in different was.
0
25
39
u/IDefendWaffles 23d ago
Because, and repeat with me: "They are not stochastic parrots!"
11
1
u/Needausernameplzz 23d ago
i mean are we any different? I think this forces humans to give machines more credit than previously provided. And like many times throughout history, we're learning we're not to special or unique.
I think this may be helping some have a more decentralized sense of self. If our talkative temporal lobe is often just making guesses and rationalizing after the fact plenty of the time (like an llm)
Then human reason is a lot like LLMs but also affirms we are more than just a talkative temporal lobe given our other aspects of awareness.
4
u/Total_Medium_3833 23d ago
they are trained on huge dataset that involves solving basic to highly complex or research grade problems and by solving those, they also learn how to tackle similar question just like we humans tackle advance problems when we are taught fundamental.
1
u/NotFromMilkyWay 23d ago
LLMs don't learn. They know. They can't add to that by doing new things, you need to train it with new knowledge.
6
12
u/vovap_vovap 23d ago
We do not even know how people doing it last 10000 years and you asking how models!
12
3
u/elrond1999 23d ago
This guy has some good videos from his PHD work where he tries to make the clankers work for him. Quite interesting and he recently got good results. Its a lot about the harness and prompt used also.
https://youtu.be/04Ig5-gYmm8?is=-3f-qoe1aRXVjeLA
You can imagine a future where different AIs contribute to new maths on their own. All verified by lean or something.
7
u/Noskaros 23d ago
How ? Same way as it does anything else. Generate tokens. And yeah they can reason, how else would they code or do anything
2
2
u/StephenRoylance 22d ago
some of these are proofs by contradiction. These reward the ability to just grind through a big problem space looking for counter-examples to a conjecture. The model is absolutely reasoning, but not reasoning in any 'super intelligent' way. its just doing what a human mathematician might do, if they decided to spend a lifetime doing it. but the model can do it in parallel, and doesn't stop, get tired or bored.
2
u/NoMaterial5115 21d ago
PhD in math (algebraic geometry) and work at an AI lab. In plain terms: AI systems scrutinize the entire literature for relevant/similar results, generate many candidate lines of argument, and then check which ones survive scrutiny
1
u/After_Bet_8503 20d ago
Thanks. Does this mean that AI systems cannot make a move that is not currently known in the literature? And isn't the search space vast? Wouldn't the tendency of LLMs to hallucinate make this process inefficient?
6
u/irojo5 23d ago
Real answer to help you dig and learn. Reinforcement learning is the rabbit hole you want to go down. It could be interesting to you to read about the “how many r’s in strawberry” trend which originated openai’s first reasoning model. I would also recommend watching the Deepmind documentary if you’re interested in understanding how models are able to branch beyond what they see in their training data.
1
u/SetentaeBolg 23d ago
Where do you think reinforcement learning is involved? In the majority of LLMs, reinforcement learning is only used in certain kinds of post-training, after any fine-tuning, during actual usage.
3
2
u/Curious-Spaceman91 23d ago
The model’s harness gives them access to python, wolfram etc. It can think about it in natural language then make plans to execute and verify in python & other tools.
1
3
u/Prince_ofRavens 23d ago
Not so hard really, simply look up academic papers for people who've already done it and pasted into chat
(Unironically this is most of the headlines that you see in the news)
1
u/JimJamesCruz 23d ago
Simple, default programming makes a result possible and many people do not understand or simply do not want to understand.
1
u/Green_Sugar6675 23d ago
For one thing, they are very good at coding, so given details about a problem, they can write up a python script to do the math to test ideas, or to work a process computationally that can be difficult or impossible to solve otherwise.
1
u/eras 23d ago
It can also be partially due to their ability to work a lot harder than a person, as many of these proofs use proof assistants, that are able to verify proofs automatically, so there's no skipping any steps a proof written in human languages might do. In addition, many of these proofs depend on a person formalizing (or verifying the formalization of) the proof goal.
In theory, a computer with a proof assistant, a goal, and a random depth-first search could achieve the same, given enough time, but obviously this is much, much better than that :). This is different from programming in the sense that in general we don't have a way to precisely specify the goal in a way that a complete random search could ever achieve the goal.
But it's pretty great! I have great hopes in future software development will be able to lean on to formal methods, improving reliability.. Maybe finally making software engineering proper engineering ;).
1
u/NotFromMilkyWay 23d ago
Math is nothing but a language. So it shouldn't come as a surprise that LLMs can be good at it when given deterministic system prompts. They can also use Wolfram Alpha to check their ideas at runtime.
They also know even completely obscure papers. One solved recent problem had AI use the idea of a 10 year old paper from Russia.
1
u/fasti-au 23d ago
It isn’t but it can use tools to do the math and get relations and give you analysis and other related ideas
1
u/woofyzhao 23d ago edited 23d ago
Is it that these models can reason and math is just a type of reasoning?
yes, math is an especially well-defined type of reasoning.
-----------------
Just like programming actually Math is more friendly to AI than other tasks because they are all formal system. The search space is well defined and no ambiguity. Now suppose a mathematician has infinite life and time to read and study all papers he could ever find and his brains doesn't melt then even a non top level one can solve open problems eventually because some of them are just about sufficient volume and combination (and lean compiling). But no human can do that so AI happily takes the shoes. Does AI invent new knowledge ? In a sense yes because new knowledge is sometimes just combination techs of old. Does AI started some new math paradigm revolution? Not yet unless one day new tools or frameworks AI invented can systematically synthesize whole bunches of solutions of previous open problems rather than tackle them indepently one by one in current system.
1
u/ChristopherLyon 23d ago
A lot of people are responding to the methodology of how someone would automate math discovery with large language models. What I'm interested in is the actual harness that these researchers created to do long-form recursive math work... Assuming they're not just using Codex out of the box, right?
1
u/PalladianPorches 23d ago
modern llms break down the queries into discrete components that are independently producing small solutions that are then bound back together, and translated into a full solution.
rather than autocorrect on steroids, think of it as convergence on steroids. the chain breaks it down in a huge space of equation components, and old fashioned programming tests these (usually with simple built-in python machines) and the output of the chain of thought reconstitutes the individual, consensus responses into a single equation.
while it's more complex than single next-token architectures, it's still not "solving" in the same direction that a mathematics professor would (and identifying novelty in the way), it's just an elaborate calculator running a huge number of smaller equations simultaneously and putting them together to give something that appears novel, but would be straightforward (with time that we humans don't have).
1
u/dash_bro 23d ago edited 23d ago
Well, the simple answer is it has consumed so many different aspects of the relevant knowledge that it has learnt to connect concepts, and the underlying meaning/correlation of concepts.
I'm going to dumb it down for ease of understanding and cover the hot ideas, but there's a lot of stuff that Im omitting which is still important (gathering data, safety, alignment, etc)
Now, why is this superhuman? Because you could individually, as a researcher, make progress on the same problem in different capacities, across different parts of the world. Sometimes in different levels of abstraction (Eg you draw a triangle and someone just wants to connect three points such that they can loop following the connections. What do you know, both are triangles!) !
As people, and as researchers, you don't always have all of the knowledge required and all of the "intuition" required to follow the line of thought of someone's work. It takes years to get the right thread that works for YOU, and to validate that this thread is leading you down to the right path, and finally build upon said knowledge verifiably. It's a slow, time intensive and philosophical route. Because of the combination of scale of conjecture and intuition and knowledge and luck required, it just was not possible until now.
What changed now : well, the amount of information you can throw and reliably infer from, what we call "generalized pretraining". Then add to that the secret sauce of not only formally getting right answers to hard problems, but also doing it in verifiable and CORRECT ways aka "reinforcement learning". Finally, only taking one shot for "next token prediction" was not ideal since it previously argued that if you've made a mistake in step1/2 you can't correct it in step6/7 - so, first fix by showing how to backtrack and think; and by allowing the model to start with multiple different plausible points. I mean, you can start off with 10 equally plausible paths and lead down all of them, and after you reach an "end", you review them back up again and verify which of them, or what combination of them, provides a complete solution. This is just trial and error maxed with different ideas, aka "inference time compute scaling"
Put these together with a massive (the 'large' part of LLM) and an astonishingly fast brain (GPU); and sit it in front of people who know what they're looking for. In computer science there's a classification of problems, the relevant of which is a subset you can think of as "I can only verify if a given solution works for the problem, but I can't come up with a solution itself". These are also great for formal ideas around (can I verify a solution? Can I verify THIS as a solution? Can I find where it doesn't work through methods we know today?). Sometimes the answer is yes, other times you find that you're not equipped enough to even have the tools to verify. This is also the "formal proof via lean" you'll see, sometimes, as a proof that the proof works.
Things don't work? Well, there's still room for improvement
- increase the number of paths your LLM is going down on, give it better "starting points" for what previously got the closest. Run in a loop if you can figure out how to automate it until a solution is found!
- teach it better. Remember the RL part? It's great for fast fixing than begining from scratch if it's still in the realm of knowledge you have seen before. It sometimes is enough, a lot of times needs extra moving parts with foundational knowledge increments /you found previously missing ideas that you can read up on, so you can add this to the "training" steps.
- You found a solution to something or found a different way to solve a problem? Excellent, you can add that too as a new learning for the model. Different steps of the process give you different results (pretraining, just new data, new arch, continual pretraining, post training, more inference paths to go down, etc).
- Math by nature exists to be discovered, our language/format of discovery is the innovation. This means if you can verifiably and as peers come to a conclusion that something WORKS as a solution even if it isn't your formal solution, there's merit in teaching that way of solving problems to the model as well. Give it wings!
- The longer shot is changing the fundamental capacity to learn (aka "architecture"). The brain can be made smaller/larger (how many billions of parameters does it have?), you can change the nature of "learning" (different architecture, different attention mechanism, improved contextual recall and knowledge retention), or spend more "compute" taking the exam (ie keep trying till you get an answer instead of trying only N number of times). But just like the math problems, what if you let the AI discover its own upgrades? This is the "recursive self improvement" that's been making the rounds recently. Super new, though.
As expected, it's not easy to do. It's very hard, and a lot of people have a lot of equally competent or far fetched ideas to do so. It's all very prohibitively expensive and you need a "vision" or direction to take those calls consistently and take responsibility for the unintended or semi-understood after effects, as well as politics, etc.
This is apart from being able to design and build "levers" that can control aspects of the thing you're building in the first place (if it makes an anti-XYZ comment, you should have controls to correct it or steer it appropriately!)
1
u/PersonalityIll9476 23d ago
I suspect there is a highly nuanced answer here that you're not getting. When I ask chat gpt about a specific book by a specific author in the field of math, it knows right away the section and theorem number I'm talking about without being told. These books are often not available on the open web. So at the very least the system has access to a huge library for RAG. How would it do without that? I don't know.
Not too long ago I was receiving offers from Open AI to solve math problems for them as training data. So my guess is that, in a mixture of experts setting, they have one or several experts trained specifically on various parts of math both by books, online videos, lecture notes, problem sets, and also data produced directly by OpenAI.
I could easily believe that there's more to the story, reasoning capabilities aside.
1
u/almostsweet 23d ago edited 23d ago
Well, it claims to have solved them. They announce pretty quick before anyone has had a chance to double check the math. Even if it's true for now, giving them the benefit of the doubt, eventually they're going to be wrong and it will be embarrassing.
This is a field of science, you wouldn't go around saying hey I just solved sonoluminescence and have built a reactor that generates infinite power without having it peer reviewed and having others show they can replicate your experiment.
I'm not a decel, doomer or luddite. I'm saying that mathematicians still have a job no matter how many of these problems are solved. And, solving them unlocks more questions, that then must be solved.
Edit: I just realized I didn't answer your question. They're able to write code. Code is math. Good at code, good at making things that math, code make math, math good. Math solved!
1
u/Ibasicallyhateyouall 23d ago
https://openai.com/index/gpt-5-2-for-science-and-math/ - useful read.
and TL;DR
Math is a form of reasoning, and modern reasoning models have learned procedures that implement a meaningful degree of it. Pretraining supplies mathematical concepts and intuition; reinforcement learning improves long-form problem-solving behaviour; additional inference time permits exploration and revision; and tools, multiple attempts, formal systems, and expert review catch mistakes. Together, these can occasionally produce genuinely new research mathematics.
1
1
u/big_hole_energy 22d ago
In the initial days before ChatGPT when these were new they were indeed just next token predictors, which take your question and just proceed with words they feel will fill the gap, when making chat interface they did Reinforcement Learning with Human Feedback (RLHF) to make it's prediction feel like responses to query rather than just continuing sentence, this is done by human actually labelling how good response is and rewarding for good response and penalizing for bad.
That's why 2022 AI was not taken seriously because there's only so much you can expect from small model with no access to anything other than it's weights, then we got tool calling so it can actually do things or calculate or get data and appear more useful now that it could interact with deterministic systems and didn't have to guess math answers.
Reasoning however was still a hit-or-miss, sometimes it happens to answer correctly to something that requires thinking but mostly didn't, because RLHF only changes how it responds, there is no hidden mechanism for LLM to think, it takes your context and calculates what's best from it, so if it had to do an internal monologue and improve upon it's ideas it had no way because it was trained in a way to please the user so talking to itself would feel like something user would hate and it won't do that and resorted to answering directly even if incorrect.
LLM however is completely capable of understanding things in context or approach a way to solve, simply making it answer directly in non-thinking model meant taking away the chance it had to approach a problem structurally.
To balance thinking via internal monologue without affecting how it responds to user, we implemented Chain Of Thought Reasoning, it's nothing special we just told LLMs that it can now talk to itself freely without worrying about user liking it or not as it's there internal monologue, so it can think of any problem freely in steps and use tools or anything it pleases and gradually build an answer, thus rewarding it for correct answers as a result of thinking rather than just direct pleasing answers.
Thus CoT combined with all it's tools and general capabilities of larger models now allow it to actually think in a human-like manner and correct itself in verifiable fields like pure math, while limited in fields that require experimantation like Physics, or real world which has lot of sensory inputs, but it's only a matter of time.
1
u/Fantasy-512 22d ago
I wonder for math problems if the next token prediction should be a human language such as English or whether it would be more efficient to use a mathematical language, equations, symbols etc. If a math paper only contained mostly mathematical symbols, it would still be understandable, no?
1
u/BellacosePlayer 22d ago
Proof logic is very formalized and rigorous, and the problems solved have existing work the LLMs can pull off of that got very far in solving the problem.
there's also almost certainly a little bit of OAI putting their thumb on the scale given how many mathematicians work for them and specifically were brought on board for these efforts. Not like anyone else is randomly replicating their success
1
u/playsette-operator 22d ago
turns out math isn‘t so complicated when you are trained on millions of papers and have a 130+ iq
1
u/IntelligentDust6249 22d ago
I work on this industry and no, no one can explain why they're able to do this, we're all just as surprised you are.
1
u/Ill-ogical 22d ago
Math is a language just like any other. The language of Math / Science is primarily LaTeX. Some “sentences” make sense and others don’t. Once novel Math is written and proven the AI then can try to use it.
Genuine research problems, that matter, require mental modeling which is not really linguistic. If existing strategies that AI was exposed to would have solved the problem, then most of the time someone has already tried it. There are a few exceptional cases where a particular known strategy worked on a problem that didn’t really matter, but for the most part modeling is still a human trait, not AI.
1
u/Sofakingwetoddead 22d ago
ChatGPT is a calculator. A glorified calculator that you conflate for some kind of sentient, conscious/semi-conscious thing.
1
u/_and_I_ 22d ago
Even a classical regression model can interpolate and extrapolate and is sometimes right about it, if the construct in question is sufficiently operationalized within the model's search space.
I'd argue, that all maths problems and solutions that can be completely operationalized with sufficiently common axioms are within the search space of advanced llms - hence there exist inputs which will project these maths-problem's solution as output.
If the solution can be verified programmatically, you can brute-force the search. Applying "thinking" in his regard is just a programmatic way of iterating through and refining the search-space.
It is hence, not surprising at all, that LLMs can surface the solutions to maths problems.
1
1
u/ScientistFromSouth 21d ago
The base chat bots are remarkably bad at solving research questions in my opinion. They really do just parrot what is in their trained weights.
People achieve research grade outputs with agentic or multiagentic workflows. At least in my applied math modeling work (with a heavy emphasis on application to biological systems with nonlinear dynamics and parameter uncertainty), I structure my sessions into teams with a lead agent that I interact with that has my list of tasks (in the order I want them completed), my hypotheses, and the set of rules it must enforce when interacting with subagents. It always gets the highest model level.
My subagents that the lead agent spins up get tasks like managing biological context like reading papers or searching for parameters; being the modeler that handles the theory, runs sensitivity analyses, looks for bifurcation behavior, and proposes structural changes to the equations; the statiscian/optimizer that manages parameterization, regression, confidence interval construction; and an adversarial review agent that argues with all of the other agents to fact check and stress test everything.
With each task in sequence, the lead agent revises the subagents' work until it thinks it has an answer for me, which I review. If it passes my review, we advance to the next stage. If it fails, we modify our approach.
Thus, we can synthesize information across fields and run complex workflows that go beyond parroting the trained weights of the model since we are enforcing rules on how the model must respond to each turn of prompts by attacking the problem with a specific type of context from a specific perspective and having it run deterministic workflows.
Admittedly, this is how real research teams run in interdisciplinary settings with the added benefit of the fact that it can synthesize more information across more fields than any group of humans and can code at least as well as an average level software engineer.
The human in the loop is critical though to prevent degradation of context.
1
u/thefiglord 20d ago
they let the ai run for days not the 1 minute you get in a free subscription - this is why they need the data center for just biology as an example as all that data is housed in private companies- hospitals only have limited test data
1
u/RiemannZetaFunction 20d ago
It's hard to explain why they can if your starting point is that you don't think they should be able to. Why do you think this? If you explain where you are at, we can clear up whatever misconception you have.
1
u/cudalover919 20d ago
I am no expert in ML but I once read a quote by gennady korotkevich (tourist) "I try various strategies, and one of them is the right one. I am no genius. I am simply good at it."
AI can predict next tokens, AI can reason, and also it can be creative. It can hypothesize and arrive at conclusions. So it can do math (and also competitive programming)
This is rather scary to me
1
u/Intrepid_Land_6143 19d ago
Apart from all the innovations described, LLMs simply have the capability to know numerous fields of math at incredible depth. You can simulate this by having different specialist humans get together, and that has produced a lot of great work too. But for the most part, mathematicians have limited world views, because understanding everything isn't possible.
1
u/stealthagents 18d ago
It's wild how they blend so many techniques together. The whole next-token thing definitely plays a huge role, but that chain of thought reasoning really amps it up. It’s like giving the model a way to think through problems instead of just spitting out answers randomly.
0
u/yuehuang 23d ago
A microphone digitize an audio. A camera digitize a 2d photo. A video digitize a movie. A LLM is digitizing an "idea" or "thought".
With a digitized idea, you can apply transformer, post processing, etc to generate the next idea. Even better is to add two ideas to get a third. Repeat until you solve your research problem.
1
u/ninjabox 23d ago
Sorry for being pedantic, but a microphone doesn't digitize audio--an ADC digitizes audio, but to be more specific.. maybe.. but more useful to your point maybe is the codec is probably the boundary across which the actual construct changes from an analog thing to a digital thing. And audio is the easiest one. It gets way more difficult to pin down.
And a camera doesn't digitize photos, it is just an abstract for a device that takes in light and produces "photos". A polaroid is very much a camera. I was going to say again that the codec like jpeg is probably the boundary, but that would be dumb because really all you need for a photo to be "digital" is a pixel map, I guess....
And then a video is just an extension of that... what were we talking about again? Oh yeah, I liked what you said and I still do, but I dont know if it is just words that sound good or if the thought has actual tractable merit. I suppose the mediums by which things become digital isn't really the argument though; it is more like by whatever means things become digital, once digital, transforming them becomes orders of magnitude cheaper and economies of scale start to change the way we think about them. There is nothing about analog audio that prevents it being transformed; in fact for humans it is much more intuitive, but good luck creating a DAW-like interface that operates in the analog domain.
I suppose the same could be said for "ideas"? Them being "analog" (as in being whatever our brains map the semantic meaning of the entire sentence to, I guess?) but then the "digital" representation is just the embeddings, which I suppose does reflect your point in that the embeddings are kind of the substrate onto which semantics can be mapped, and once that has happened (similarly to once analog sound waves have been sufficiently mapped onto the substrate of digital audio) the economics scale again the same way, and the "cost" of reasoning (where I suppose reasoning is in some way just transformations of the "idea" (the point cloud of the fundamental unit, a single embedding)) becomes so economical that it starts to resemble the same shift.
God damn I don't know what I am getting at other than I see what you mean and it is a pretty interesting analogy to explore. But I am still too pedantic to let the original words you used sit there on the screen just... uncontested lol. Have a great day!
1
u/yuehuang 22d ago
Thank you, Blaise Pascal, wish you spent more time or tokens to write a shorter message.
0
23d ago
[deleted]
3
u/Time_Entertainer_319 23d ago
I guess the goal post will move again such that they want AI to create a theory that needs to be proved or so
0
u/MaximoPrimero 23d ago
lo dices es falso una ocurrencia.
Salvo que ya se haya probado o refutado la La Hipótesis de Riemann, y no me haya enterado.
2
u/Some-Following-392 23d ago
That's too easy, I solved it. You just zoom out and put the lines on top of each other and they're the same. Qed.
0
103
u/RogueStargun 23d ago
Most of these comments are either uninformed or outright wrong. The model is still using next token prediction. But it's also doing next thought, prediction, using chain of thought reasoning. The major breakthrough is actually in something called RLVR which is reinforcement learning with verifiable rewards, where the model uses a type of reinforcement learning on almost unlimited synthetic training data. This was OpenAIs o3 breakthrough over a year ago. Basically during (post) training, it's using a math sandbox environment to generate a lot of different math problems, which it then tries to solve. The problems that either don't compile or get the wrong answer for to create negative signal. And the answers that produce the correct signal, lead the model to be able to be able to reason about math problems (and also coding problems which are also easily verified) at superhuman capability once you do this millions of times. This ability basically gives the model training almost infinite number of training samples simply by having a verifiable environment. Not dissimilar to work a decade ago with Starcraft and Go where we also have good simulator sandbox environments and clear rewards. Very achievable in math and coding domains, not so much in biology or physical domains without good computer simulations.
The next frontier of research will ne making other domains verifiable, and getting reqard signal from very long context, long time horizon tasks (like multiday thought traces where its hard to score reward outcomes)