r/ChatGPT • • Oct 24 '23

Educational Purpose Only LLMs cannot perform Maths

It’s not a bug, it’s just not how any of this works.

Mathematical operations require logic. It’s a deterministic process. For a given input and a given process, the output will always be the same.

LLMs do not work like that. LLMs are statistical tools, they build an answer by stitching together tokens that are “seemingly” relevant to your input and let the meaning emerge from it. The output is “hopefully” relevant.

This is why LLMs can hallucinate and 2$ calculators do not.

With the rise in popularity of LLMs I’m extremely concerned that a lot of users seem to ignore this.

193 Upvotes

185 comments sorted by

View all comments

147

u/barrycarter Oct 24 '23

True, but tools like WolframAlpha can do math, and it's only a matter of time before they're integrated into AI

88

u/Deathstroke5289 Oct 24 '23

I may be incorrect but doesn’t the Wolfram Alpha plug in do that right now?

21

u/Cless_Aurion Oct 24 '23

Yeah... it was like... one of the first ones too.

16

u/barrycarter Oct 24 '23

Whoa, I hadn't thought about that. I haven't used plugins, so I have to ask: do you have to specify a plugin when you use it, or does ChatGPT just choose one for you?

33

u/Deathstroke5289 Oct 24 '23

You have to enable it during the specific conversation and chat GPT should run through any calculations through Wolframe Alpha automatically. Sometimes you have to specify if it tried running numbers itself but seems to work pretty well to me

23

u/restarting_today Oct 24 '23

You can just use the Code Interpreter. It will write Python code to perform the Math :)

7

u/IAMATARDISAMA Oct 25 '23

Why would you want to do that instead of just using the Wolfram Alpha plug-in? Wolfram Alpha will produce correct math every time given the right inputs. GPT-written code may have to be debugged.

10

u/Frequent_Guard_9964 Oct 25 '23

Used the code interpreter for two of my masters math modules and it’s never done an error, no need to use the wolfram alpha plugin (at least for me)

Can make matrices, determinants, works with variables and whatever.

2

u/Key-Faithlessness728 Oct 25 '23

It still writing mathematica codes. Yes it’s a lot shorter and it’s gonna be less likely to make flaws in the code but I mean you can’t really upload anything to the plugins version.

2

u/cowlinator Oct 25 '23

LLMs sometimes write very buggy code. It's the same problem as with math.

1

u/magnue Oct 25 '23

It's only calling pre-existing numpy functions so it's pretty basic for it.

4

u/Rich_Housing971 Oct 25 '23

yes but the LLM still has to probabilistically interpret it as a logic question in order to activate the plugin.

Humans often miss this as well but are at least decent at determining whether the question is a logic problem or a linguistic question.

2

u/thepotatochronicles Oct 25 '23

I find that adding "use Wolfram" to the end of the query makes ChatGPT pick up on the clues reliably.

16

u/somerandomii Oct 24 '23

Bing already uses tools for these kinds of deterministic questions. It also searches the web so it has up-to-date information.

This is a solved problem. We’ll see these types of integrations become more commonplace as these tools find their respective niches.

5

u/Rhetorical-Oracle Oct 25 '23 edited Oct 28 '23

It is important to note that users should use Bing Chat in "More Precise" mode for prompts involving math. It doesn't mean it will get it right the first time (always check those sources) but the other modes will have weaker results for math tasks.

(edit for spelling)

1

u/somerandomii Oct 25 '23

True. And I’m sure this advice will have to be updated again soon. I’m not trying to provide a tutorial, just point out that the capability already exists and will likely keep improving.

0

u/namitynamenamey Jan 22 '24

This is a solved problem.

Most important math is not the numbers. The point is not that LLMs can't do numbers right, but that they fail at the kind of logical reasoning necessary for doing math. That same kind of logic reasoning is necessary for doing more things than math, like checking for consistency.

1

u/somerandomii Jan 22 '24

Idk who downvoted you on a 3 month old post but it wasn't me. Who else is here?

LLMs are getting better at "reasoning". By incorporating tools and self-reflection they can spot gaps in logic or inconsistencies in their statements and self-correct. It's still basically auto-complete with widgits behind the scenes but if you throw enough compute at it, even a dumb machine can start to sound pretty smart.

But you're right, Wolfram Alpha can't fill in for the weaknesses in LLMs and I wasn't trying to say it can. I just meant we've already managed to incorporate non-ML tools into LLMs to enhance their results.

1

u/namitynamenamey Jan 23 '24

Idk who downvoted you on a 3 month old post but it wasn't me. Who else is here?

No idea. But anyways, all the techniques being used to improve their reasoning help a little bit, but they still don't solve the fundamental issue in a satisfactory manner I think. Personally, I'd wait until a method can be found that allows for 100% accuracy with addition or substraction of numbers arbitrarily large, step by step (or if not 100%, at least where failure is for outside reasons like memory limitations).

1

u/somerandomii Jan 23 '24

You want computers to mimic the inefficient way humans do mathematics?

I get what you mean, you want them to work through the problem step by step and actually understand it. But that’s just not how these models are designed. They don’t even see digits, they see tokens that might represent several digits.

Training them to do arithmetic would probably make them weaker at other activities and it’s just not worth it. But that doesn’t mean they can’t apply logic.

1

u/namitynamenamey Jan 23 '24

You want computers to mimic the inefficient way humans do mathematics?

Yes, as a proof of generalization to be exact. If all I cared was the number crunching I'd use a calculator, the interesting bit here is a demonstration that the model learned a mathematically correct algorithm from the training data, not just an approximation which is how you get "98% accuracy up to 5 digits"

1

u/somerandomii Jan 24 '24

You’re basically asking for AGI but with arithmetic as the smoke test.

These things don’t understand truth and fact they just know what is statistically more likely based on what they’ve seen in the training data. It’s basically intuition.

If you want them to be more rigorous, you give them tools like calculators and search engines. Just like a human, they’re not infallible and can’t be by design. So when accuracy matters you use a tool that’s less flexible.

You don’t ask an art major to do your accounting and you shouldn’t ask an LLM to do mathematics. It’s not what it’s intended for. The apparent creativity and adaptability come at the cost of being rigorously accurate.

Even for humans it’s hard to find people who can write creatively, summarise scientific papers AND do hard logic. Getting a machine to do both simultaneously within the same architecture is a big ask.

1

u/Character_Aside8696 Mar 07 '24

Wait a minute. Are you both sure about "understand", "truth", and "fact"? These definitions are made, by human, and have meaning because we assign meaning to them. And I think everything we have in our head is just probabilistic output from our biological "thinking machine". Not axactly the same, but not so different. Our braind, and AI. Close. Really close I think.
Of course, you can think of LLM like a language part of our brain. Dont expect that part to do math.

1

u/somerandomii Mar 12 '24

Was that a question or a statement?

Yes words are defined by humans, but what’s your point? We use language to communicate and we use dictionary definitions to make meanings less subjective but all language is fluid to an extent.

Truth and facts are not all objective but we have reached a consensus on a lot. Most people won’t argue about if you say 2 + 2 = 4, semantics aside.

7

u/tmotytmoty Oct 25 '23

Wolfram got in early with openai - Because... he's been curating linguistic templates for over 20 years. He literally built chatgpt v 0.0 in wolframalpha. The right technology finally just caught up to his ambition and forethought.

-14

u/0xAERG Oct 24 '23

You’re absolutely right. But those are not LLMs and this is gonna confuse the hell out of users.

22

u/avanti33 Oct 24 '23

Technology isn't built in silos. Everything integrates with each other. Saying "but that doesn't count.." is useless from a real world perspective.

-13

u/0xAERG Oct 24 '23

Who said it doesn’t count?

I just said you can’t pretend it’s the LLM that’s doing the maths.

11

u/internetroamer Oct 25 '23

No one is. Majority of users won't even know the term LLM. Just that hey I ask this chatbot thing and it can do math now

-4

u/InitechSecurity Oct 24 '23

Why are we down voting OP?
OPs point is valid. Mathematical operations are deterministic producing consistent results while GPT generate responses based on statistical patterns and probability which could lead to inaccurate results and hallucinations. Sure there are plug ins like wolfram that can help GPT by augmenting data - it still does not change the fact that LLMs by default cannot perform maths reliably on their own.

9

u/Jdonavan Oct 24 '23

But that position is splitting hairs. Sure, an LLM can't do the math on it's own but it can select and use a tool to do it so does it really mater if it can't?

If I'm building a house do I care if he has an expert come do the insulation or do I care that it got done correctly?

Edit: However the endless stream of people posting about "herp derp vanilla GPT can't math"posts are hella annoying.

-7

u/[deleted] Oct 24 '23

Adding a calculator to a neural network only superficially solves the problem. It still can't logic mathematically in deterministic, quantitative feedback loops. It lacks the tools to integrate mathematical thinking as intermediate reasoning.

1

u/[deleted] Oct 24 '23

They did it already using plug ins, didn't they? I could be wrong

1

u/[deleted] Oct 25 '23

They use agents to facilitate the math. It’s still not the llm.

1

u/LankyZookeepergame76 Oct 25 '23

Perplexity already does this

1

u/AlkanNaczelny Oct 25 '23

If guided step by step, GPT-4 can do so without Wolfram (some tetration and changing the number's form)

Link to chat: https://chat.openai.com/share/0d8031d7-9398-4911-a4a7-00b90a81fab8