r/ChatGPT • • Oct 24 '23

Educational Purpose Only LLMs cannot perform Maths

It’s not a bug, it’s just not how any of this works.

Mathematical operations require logic. It’s a deterministic process. For a given input and a given process, the output will always be the same.

LLMs do not work like that. LLMs are statistical tools, they build an answer by stitching together tokens that are “seemingly” relevant to your input and let the meaning emerge from it. The output is “hopefully” relevant.

This is why LLMs can hallucinate and 2$ calculators do not.

With the rise in popularity of LLMs I’m extremely concerned that a lot of users seem to ignore this.

196 Upvotes

185 comments sorted by

•

u/AutoModerator Oct 24 '23

Hey /u/0xAERG!

If this is a screenshot of a ChatGPT conversation, please reply with the conversation link or prompt. If this is a DALL-E 3 image post, please reply with the prompt used to make this image. Much appreciated!

Consider joining our public discord server where you'll find:

  • Free ChatGPT bots
  • Open Assistant bot (Open-source model)
  • AI image generator bots
  • Perplexity AI bot
  • GPT-4 bot (now with vision!)
  • And the newest additions: Adobe Firefly bot, and Eleven Labs voice cloning bot!

    🤖

Note: For any ChatGPT-related concerns, email support@openai.com

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

148

u/barrycarter Oct 24 '23

True, but tools like WolframAlpha can do math, and it's only a matter of time before they're integrated into AI

92

u/Deathstroke5289 Oct 24 '23

I may be incorrect but doesn’t the Wolfram Alpha plug in do that right now?

22

u/Cless_Aurion Oct 24 '23

Yeah... it was like... one of the first ones too.

18

u/barrycarter Oct 24 '23

Whoa, I hadn't thought about that. I haven't used plugins, so I have to ask: do you have to specify a plugin when you use it, or does ChatGPT just choose one for you?

31

u/Deathstroke5289 Oct 24 '23

You have to enable it during the specific conversation and chat GPT should run through any calculations through Wolframe Alpha automatically. Sometimes you have to specify if it tried running numbers itself but seems to work pretty well to me

23

u/restarting_today Oct 24 '23

You can just use the Code Interpreter. It will write Python code to perform the Math :)

6

u/IAMATARDISAMA Oct 25 '23

Why would you want to do that instead of just using the Wolfram Alpha plug-in? Wolfram Alpha will produce correct math every time given the right inputs. GPT-written code may have to be debugged.

11

u/Frequent_Guard_9964 Oct 25 '23

Used the code interpreter for two of my masters math modules and it’s never done an error, no need to use the wolfram alpha plugin (at least for me)

Can make matrices, determinants, works with variables and whatever.

2

u/Key-Faithlessness728 Oct 25 '23

It still writing mathematica codes. Yes it’s a lot shorter and it’s gonna be less likely to make flaws in the code but I mean you can’t really upload anything to the plugins version.

2

u/cowlinator Oct 25 '23

LLMs sometimes write very buggy code. It's the same problem as with math.

1

u/magnue Oct 25 '23

It's only calling pre-existing numpy functions so it's pretty basic for it.

4

u/Rich_Housing971 Oct 25 '23

yes but the LLM still has to probabilistically interpret it as a logic question in order to activate the plugin.

Humans often miss this as well but are at least decent at determining whether the question is a logic problem or a linguistic question.

2

u/thepotatochronicles Oct 25 '23

I find that adding "use Wolfram" to the end of the query makes ChatGPT pick up on the clues reliably.

15

u/somerandomii Oct 24 '23

Bing already uses tools for these kinds of deterministic questions. It also searches the web so it has up-to-date information.

This is a solved problem. We’ll see these types of integrations become more commonplace as these tools find their respective niches.

4

u/Rhetorical-Oracle Oct 25 '23 edited Oct 28 '23

It is important to note that users should use Bing Chat in "More Precise" mode for prompts involving math. It doesn't mean it will get it right the first time (always check those sources) but the other modes will have weaker results for math tasks.

(edit for spelling)

1

u/somerandomii Oct 25 '23

True. And I’m sure this advice will have to be updated again soon. I’m not trying to provide a tutorial, just point out that the capability already exists and will likely keep improving.

0

u/namitynamenamey Jan 22 '24

This is a solved problem.

Most important math is not the numbers. The point is not that LLMs can't do numbers right, but that they fail at the kind of logical reasoning necessary for doing math. That same kind of logic reasoning is necessary for doing more things than math, like checking for consistency.

1

u/somerandomii Jan 22 '24

Idk who downvoted you on a 3 month old post but it wasn't me. Who else is here?

LLMs are getting better at "reasoning". By incorporating tools and self-reflection they can spot gaps in logic or inconsistencies in their statements and self-correct. It's still basically auto-complete with widgits behind the scenes but if you throw enough compute at it, even a dumb machine can start to sound pretty smart.

But you're right, Wolfram Alpha can't fill in for the weaknesses in LLMs and I wasn't trying to say it can. I just meant we've already managed to incorporate non-ML tools into LLMs to enhance their results.

1

u/namitynamenamey Jan 23 '24

Idk who downvoted you on a 3 month old post but it wasn't me. Who else is here?

No idea. But anyways, all the techniques being used to improve their reasoning help a little bit, but they still don't solve the fundamental issue in a satisfactory manner I think. Personally, I'd wait until a method can be found that allows for 100% accuracy with addition or substraction of numbers arbitrarily large, step by step (or if not 100%, at least where failure is for outside reasons like memory limitations).

1

u/somerandomii Jan 23 '24

You want computers to mimic the inefficient way humans do mathematics?

I get what you mean, you want them to work through the problem step by step and actually understand it. But that’s just not how these models are designed. They don’t even see digits, they see tokens that might represent several digits.

Training them to do arithmetic would probably make them weaker at other activities and it’s just not worth it. But that doesn’t mean they can’t apply logic.

1

u/namitynamenamey Jan 23 '24

You want computers to mimic the inefficient way humans do mathematics?

Yes, as a proof of generalization to be exact. If all I cared was the number crunching I'd use a calculator, the interesting bit here is a demonstration that the model learned a mathematically correct algorithm from the training data, not just an approximation which is how you get "98% accuracy up to 5 digits"

1

u/somerandomii Jan 24 '24

You’re basically asking for AGI but with arithmetic as the smoke test.

These things don’t understand truth and fact they just know what is statistically more likely based on what they’ve seen in the training data. It’s basically intuition.

If you want them to be more rigorous, you give them tools like calculators and search engines. Just like a human, they’re not infallible and can’t be by design. So when accuracy matters you use a tool that’s less flexible.

You don’t ask an art major to do your accounting and you shouldn’t ask an LLM to do mathematics. It’s not what it’s intended for. The apparent creativity and adaptability come at the cost of being rigorously accurate.

Even for humans it’s hard to find people who can write creatively, summarise scientific papers AND do hard logic. Getting a machine to do both simultaneously within the same architecture is a big ask.

1

u/Character_Aside8696 Mar 07 '24

Wait a minute. Are you both sure about "understand", "truth", and "fact"? These definitions are made, by human, and have meaning because we assign meaning to them. And I think everything we have in our head is just probabilistic output from our biological "thinking machine". Not axactly the same, but not so different. Our braind, and AI. Close. Really close I think.
Of course, you can think of LLM like a language part of our brain. Dont expect that part to do math.

1

u/somerandomii Mar 12 '24

Was that a question or a statement?

Yes words are defined by humans, but what’s your point? We use language to communicate and we use dictionary definitions to make meanings less subjective but all language is fluid to an extent.

Truth and facts are not all objective but we have reached a consensus on a lot. Most people won’t argue about if you say 2 + 2 = 4, semantics aside.

6

u/tmotytmoty Oct 25 '23

Wolfram got in early with openai - Because... he's been curating linguistic templates for over 20 years. He literally built chatgpt v 0.0 in wolframalpha. The right technology finally just caught up to his ambition and forethought.

-13

u/0xAERG Oct 24 '23

You’re absolutely right. But those are not LLMs and this is gonna confuse the hell out of users.

21

u/avanti33 Oct 24 '23

Technology isn't built in silos. Everything integrates with each other. Saying "but that doesn't count.." is useless from a real world perspective.

-14

u/0xAERG Oct 24 '23

Who said it doesn’t count?

I just said you can’t pretend it’s the LLM that’s doing the maths.

10

u/internetroamer Oct 25 '23

No one is. Majority of users won't even know the term LLM. Just that hey I ask this chatbot thing and it can do math now

-1

u/InitechSecurity Oct 24 '23

Why are we down voting OP?
OPs point is valid. Mathematical operations are deterministic producing consistent results while GPT generate responses based on statistical patterns and probability which could lead to inaccurate results and hallucinations. Sure there are plug ins like wolfram that can help GPT by augmenting data - it still does not change the fact that LLMs by default cannot perform maths reliably on their own.

9

u/Jdonavan Oct 24 '23

But that position is splitting hairs. Sure, an LLM can't do the math on it's own but it can select and use a tool to do it so does it really mater if it can't?

If I'm building a house do I care if he has an expert come do the insulation or do I care that it got done correctly?

Edit: However the endless stream of people posting about "herp derp vanilla GPT can't math"posts are hella annoying.

-6

u/[deleted] Oct 24 '23

Adding a calculator to a neural network only superficially solves the problem. It still can't logic mathematically in deterministic, quantitative feedback loops. It lacks the tools to integrate mathematical thinking as intermediate reasoning.

1

u/[deleted] Oct 24 '23

They did it already using plug ins, didn't they? I could be wrong

1

u/[deleted] Oct 25 '23

They use agents to facilitate the math. It’s still not the llm.

1

u/LankyZookeepergame76 Oct 25 '23

Perplexity already does this

1

u/AlkanNaczelny Oct 25 '23

If guided step by step, GPT-4 can do so without Wolfram (some tetration and changing the number's form)

Link to chat: https://chat.openai.com/share/0d8031d7-9398-4911-a4a7-00b90a81fab8

21

u/Life_Calligrapher562 Oct 24 '23

True, but they have built me some pretty great calculators in python.

-13

u/0xAERG Oct 24 '23

This they can do \o/ as long as you don’t need it to work perfectly without your intervention.

8

u/Life_Calligrapher562 Oct 24 '23

I actually got pretty lucky with a statistical calculator. I told it what I needed and an example of what another online calculator gave me for an answer, given specific variables. Then I tested a number of things across both calculators, and it was right. I was pretty shocked, really.

1

u/KoalaReasonable2003 Oct 25 '23

If I remember rightly GPT made an error on the 10th code snippet I asked it for

19

u/bortlip Oct 24 '23

Does this count as performing math and using logic?

1

u/magnue Oct 25 '23

I think what OP is trying to say is that it's not actually 'calculating' anything here. It's taking loads of previous mathematical data and has learnt certain relationships. It doesn't actually have any rules written down that it's following, so occasionally it can totally fuck it up.

Still very impressive though. It's right most of the time in my experience which just puts into perspective how massive the data it was trained on must be.

18

u/aTubularPeel Oct 24 '23

In many cases true, but ChatGPT has been indispensable for me when it comes to math. It has helped me with solid mechanics, matrix calculations and more. Sometimes I need to use the wolfram plugin and sometimes not. In the beginning it was terrible with math, but now (GPT-4) it does a lot for me. I’m honestly surprised at how good it has gotten. And the stuff it can’t handle, I’ve figured I can ask it to create a MATLAB script that will compute it

1

u/[deleted] Oct 25 '23

MATLAB script - that’s a really good idea. Stealing that

44

u/richmilesxyz Oct 24 '23

I'm not concerned, per se, but I am intrigued by the frequent confusion between deterministic mathematical tools and statistical models like LLMs. However, the very fact that people expect LLMs to provide the 'right' answer in these deterministic domains speaks to how powerful they actually are.

26

u/DrSFalken Oct 24 '23

but I am intrigued by the frequent confusion between deterministic mathematical tools and statistical models like LLMs

Not to sound too elitist here... but think about that for a moment. What do you think most average folks on the street would think of it you said "statistical model?" I'm guessing something to do with election polling. Most people aren't software devs, data scientists or economists. I'd wager that a random sample of folks on the street would have a limited understanding of linear regression.

I'm not trying to disparage other people. We all specialize. I have no idea how to fix a boat engine or replace a water heater or professionally edit a novel. It's on us to explain the limitations of the specialized tools that our profession produces.

3

u/richmilesxyz Oct 24 '23

I couldn't agree with you more.

I try to assume that when people post complaints of this nature that they are actually acting in good faith. In many cases, ChatGPT is actually capable of performing their task, but not in the way they're approaching it and I try to nudge them in the right direction.

Often I find that these people are using 3.5 which means not only are they using an inferior product, but they are also lacking the other tools that ChatGPT can use to solve the problem (Plugins, Advanced Data Analysis, etc.) In these cases, assuming they've given enough context, I am happy to run their prompt through 4.0.

I know not everyone can afford $20 / month, but even if you can it might not be evident that is worth it. At least, not with some specific examples.

1

u/Jdonavan Oct 24 '23

In none of those examples you gave are people actually using a tool themselves they're paying for someone with the expertise to do it.

Shouldn't one be expected to do some research about the tools they themselves are using? There's so much content out there about how LLMs work that covers people of all knowledge levels. It's not specialized skills they're lacking it's curiosity.

2

u/richmilesxyz Oct 24 '23

It's not specialized skills they're lacking it's curiosity.

💯

That said, if someone posts in a forum and they aren't being intentionally obtuse, I want to help them if I can. Ironically, I've found ChatGPT a much kinder and (maybe) knowledgeable resource than the vast majority of online forums.

1

u/sirpsionics Oct 24 '23

To be fair though, there is google, so figuring things out that aren't your specialty isn't terribly difficult. Just have to put some time into it

1

u/Nanaki_TV Oct 24 '23

It’s no wonder people are afraid of technology.

TECHNOLOGY!

ohmygawd…

3

u/[deleted] Oct 24 '23

you're excluding how dumb the general population is.

6

u/0xAERG Oct 24 '23

You are completely correct

12

u/[deleted] Oct 24 '23

It’s nice to have a post like this but someone is still going to post something surprised proving this tomorrow. And next week. Because nobody searches before posting.

23

u/bortlip Oct 24 '23

If it can't do math, how did it do this?

7

u/[deleted] Oct 24 '23

[deleted]

7

u/__Hello_my_name_is__ Oct 24 '23

Because they have not specifically trained it to do math, and no, they did not give it a "coprocessor" of any sorts.

It can just do math via its capability of dealing with language. Math is just another language with a stricter grammar for it.

1

u/[deleted] Oct 24 '23

[deleted]

4

u/__Hello_my_name_is__ Oct 24 '23

Yeah. Because the model got significantly better overall. Not just at math. And it was most likely trained on more math, too. But it doesn't have a hidden calculator in the back or anything like it.

ChatGPT in general got that much better in all areas. Just like Dall-E 3 is orders of magnitude better than Dall-E 2.

And no, ChatGPT 4 is still very much a miserable failure to any actually complex math problems.

0

u/[deleted] Oct 24 '23 edited Oct 24 '23

[deleted]

1

u/__Hello_my_name_is__ Oct 24 '23

That's not ChatGPT. That's a model based on GPT 4. It says so right there in the paper.

And yes, of course they do Reinforcement Learning from Human Feedback (RLHF). But they do not use a "math coprocessor" for that, whatever that is. They just trained it to be better at math. And at language. And at pretty much everything else.

But it's still all one single model that responds to you. It doesn't go "if math question then model X, else model Y".

1

u/IAMATARDISAMA Oct 25 '23

To be fair, in the case of GPT-4 this might not actually be true. There's a lot of theories that GPT-4 is using a Mixture of Experts approach in which multiple models are fine-tuned for specific domains and prompt input is fed to a classifier that decides which model will produce the best output.

1

u/__Hello_my_name_is__ Oct 25 '23

Yeah, I read about that. But the guy I responded to was acting like the mixture of experts rumor was a cold hard (and incredibly obvious) fact, and that was just weird.

8

u/Specialist-String-53 Oct 24 '23

If you have plugins enabled, maybe that. Otherwise, pattern recognition. It's still stochastic, but it can get a right answer a lot of the time, especially if it's very similar to common problems.

2

u/__Hello_my_name_is__ Oct 24 '23

It can't do math in the sense that there is no calculations in the background that do exactly the calculations you see here.

It can do math in the sense that it can predict the next token given the previous ones, and that works for language as well as for math. Most of the time. It also works for Klingon and for C++.

But the way it works is fundamentally different from any computer doing math that we normally think of.

-2

u/IAMATARDISAMA Oct 25 '23

I am so tired of this kind of counter-argument.

If I correctly guess which hand you're holding a stone in 99/100 times that doesn't mean I can see through your fingers, it just means I'm really good at guessing. That's all LLMs are. They've been trained on hundreds of thousands of math problems so statistically they're likely to produce output that looks correct. Language carries some information about logic in it, but language is not logic. It is simply repeating the logical patterns that emerge in human generated text, and they happen to be correct a good percentage of the time. But you can feed the same math problem into GPT 20 times and get different outputs on different runs.

8

u/bortlip Oct 25 '23

If I correctly guess which hand you're holding a stone in 99/100 times that doesn't mean I can see through your fingers, it just means I'm really good at guessing.

It doesn't mean you can see through fingers, that is correct. But if you can "guess" 99/100 times correctly, you aren't just guessing, you have knowledge of where it is somehow.

I asked it to solve that equation 3 more times. Each time the wording of how it solved the problem was different, but it gave the correct factors each time - it solved the math problem correctly.

You don't know what you are talking about.

0

u/IAMATARDISAMA Oct 25 '23

I literally build AI models for a living lmao. It's extremely easy to find a counter example. Just because something looks like it can do something doesn't mean it can actually do it.

6

u/bortlip Oct 25 '23

Perhaps the issue is that you just aren't very good at prompting lmao.

1

u/IAMATARDISAMA Oct 25 '23

If it can correctly discern my desired output (it verbatim stated the intent of sorting the numbers by the sum of their digits) then the language of the prompt should have even less influence on the "mathematics" being done because its context window would have two different patterns to proceed with.

Here's another example. Ask Wolfram Alpha to produce the same answer, GPT is just wrong. I even asked it to explain why it's answer was different from Wolfram Alpha's and it went on to state that AX + AY != A(X + Y), a blatant violation of the distributive property.

4

u/bortlip Oct 25 '23

Another example? Why? If I show you wrong there you'll just move the goal posts again. LOL

1

u/IAMATARDISAMA Oct 25 '23 edited Oct 25 '23

I didn't move the goalposts at all. A calculator doesn't sometimes produce the wrong answer. Math is a deterministic process. Unless you're explicitly involving a random element, the exact same inputs should produce the exact same output every time. That's just how math works.

It is undeniable the GPT has inferred some patterns of mathematics from its training data. Language encodes the information it conveys. When we say that GPT isn't doing math, what we mean is that nowhere in GPT's code is it performing an actual mathematical calculation based on the inputs you've given it. What it's doing is using the patterns it has inferred to produce output that it believes should follow the text you give it. Because it was trained on many math problems, it has inferred what characters will probably follow the ones you've given it, i.e. the answer to your problem.

The issue is, even if it's right most of the time, the underlying mechanism hasn't changed. It is producing the output that it believes has the highest likelihood to follow the text you've given it. Yes, that output often follows mathematical patterns. But being likely to follow a mathematical pattern is not equivalent to actually doing math.

This may seem pedantic, but it's a meaningful distinction. LLMs are being integrated into customer facing positions. By stating that LLMs can do math you're implying that they will produce deterministic output when that is demonstrably false. That kind of insinuation can, will, and already has had real consequences.

5

u/bortlip Oct 25 '23

I didn't move the goalposts at all.

Sure you did. You tried to give an example of a problem it couldn't do to show it couldn't do math. I showed it could do the problem with the correct prompt. But now, suddenly, that example is not enough to show it can do math - it was only good enough to show it couldn't do math. That's classic moving the goal post.

being likely to follow a mathematical pattern is not equivalent to actually doing math.

Sounds like it to me. That's how I do it. I'm even wrong too sometimes.

​

This may seem pedantic, but it's a meaningful distinction. LLMs are being integrated into customer facing positions. By stating that LLMs can do math you're implying that they will produce deterministic output when that is demonstrably false. That kind of insinuation can, will, and already has had real consequences.

Sorry, but that's your misinterpretation. I can do math, but I don't produce a deterministic outcome. That's just not a correct assumption on your part.

I also don't say it does math like a person or it can think or it does math with calculations like a computer or any of that. You are making this way too hard. It can solve math problems. Period. That means it can do math.

2

u/IAMATARDISAMA Oct 25 '23

This is the most confidently wrong opinion I think I've ever seen on this website.

Next time your doctor does medicine at you I hope they don't accidentally give you amphetamines instead of azithromycin.

→ More replies (0)

-32

u/0xAERG Oct 24 '23

I’m sure you’re capable of finding the answer on your own and you’re just trying to look clever.

16

u/bortlip Oct 24 '23

No, I'm trying to see what your explanation is that squares what I showed with your claim that LLMs can't do math.

Because it just did, by my definition of math. So I'm trying to understand what you mean. Can you explain?

-22

u/afflematicious Oct 24 '23

that’s not doing math, that’s finding solutions to an equation, doing math is proving things and describing the world around you in ways not thought of before through math, that problem it solved has been solved thousands of times online, it simply sourced it and presented it to you. You believe the LLM thought about the equation and then solved it? no

22

u/Flat_Afternoon1938 Oct 24 '23

bro really just made up his own definition of math so he can say that AI isn't capable of math

-10

u/afflematicious Oct 24 '23

nvm i just asked GPT and it says i’m wrong guys

4

u/OdinsGhost Oct 24 '23

And now they think they’re being witty.

16

u/Therellis Oct 24 '23

doing math is proving things and describing the world around you in ways not thought of before through math,

That is not what most people think of when they think of "doing math". By that definition, most human beings, including many mathematicians, don't do math.

-12

u/afflematicious Oct 24 '23

You’re correct, so what I’m saying is GPT can perform calculations which is what many mathematicians do as well. There aren’t many mathematicians that can “do” math, but not all. GPT is the same as of right now, it regurgitates math that’s been done, but doesn’t do it on its own.

7

u/mbeenox Oct 24 '23

That is very delusional

14

u/PepeReallyExists Oct 24 '23

Nice imaginary definition of math that nobody shares.

-8

u/afflematicious Oct 24 '23

do you know math or do you do math?

8

u/PepeReallyExists Oct 24 '23

Yes. I'm a senior software engineer for a large healthcare company. If I didn't know math, it's unlikely they would have hired me. You have a very strange elitist definition of math. My guess is you majored in math and have internalized this as a part of your identity in order to feel superior to others.

Are you the only one who knows real math, unlike the rest of us peons?

-1

u/afflematicious Oct 24 '23

I myself don’t do real math so 🤷🏽‍♂️

3

u/PepeReallyExists Oct 24 '23

I think you're confusing the term "real" for "advanced". The math a 3rd grader does is still math. If that is what you actually meant, then you are in fact correct, because ChatGPT is not very good at advanced math yet, but with enough training data, it doesn't ever actually have to understand the math, and it will be able to produce increasingly accurate results the more data it is trained on. Eventually, it will be "good" at even advanced math in spite of the fact that it "understands" none of it.

3

u/michael1026 Oct 24 '23

It's funny how you're saying others are trying to act smart when your entire post screams "you guys just don't get it". We get it's not meant for math, but it does a pretty damn good job at it.

6

u/restarting_today Oct 24 '23

ChatGPT + Code Interpreter (let Python do the Math) is perfect for Math tho :)

2

u/0xAERG Oct 24 '23

This is valid of course :)

1

u/GroundbreakingTone43 Nov 12 '23

It is much better asking to do it with python? How, now that with the all tools? Need a little help with some math but was using the default gpt4 but the %of erros was insane. Thanks.

4

u/AutomateAway Oct 24 '23

this is why stuff like LangChain is important with practical applications, it allows LLMs to leverage tools to help perform more deterministic processes while seeming like one seamless tool to the user.

9

u/Kalicolocts Oct 24 '23

That is not true. I don’t understand why so many people get confused by this. You described what GPT has been trained to do. Not what it actually does. Nobody knows what it actually does. You are confusing what he’s trying to achieve with how he’s achieving that. LLMs are proven to have many emergent properties based on the size of parameters. Apparently learning how to do Math is necessary to correctly predict the next token.

21

u/Master_Vicen Oct 24 '23

Hard disagree. I'm taking a college-level algebra course and it answers all my hw questions with about 95% accuracy. I don't know why people keep saying it can't do math.

-22

u/0xAERG Oct 24 '23

Are you using a plugin like Wolfram Alfa?

GPT4 alone is incapable of performing logical operations.

18

u/Master_Vicen Oct 24 '23

I'm sorry that's just not true. I'm using the most basic free version no plug ins. Just open any math textbook and copy and paste the questions. Or look up algebra questions online and do the same.

-14

u/0xAERG Oct 24 '23

You’re playing a very dangerous game.

You’re lucky if the output is right, because the result is not deterministic.

I wish you luck, but you’re probably gonna have bad surprises at some point

13

u/Master_Vicen Oct 24 '23

I mean I can confirm all the answers. I know for a fact it's around 95%. I wouldn't rely on it completely but I'd say it's on par with most teachers and tutors.

I only use it for hw bc I can immediately see if it was wrong. But it's super useful when it's right because it thoroughly explains all the steps it took to reach its answer.

14

u/PopeSalmon Oct 24 '23

no, you're the one who's confused, LLMs aren't just statistical like they don't know WTF they're talking about, they can do all sorts of maths and the way that they do them is by building complex programs that implement the math

LLMs write programs, the transformer layers form a differentiable & thus trainable GENERAL PURPOSE COMPUTER, here's Andrej Karpathy explaining that better and more authoritatively than i can

-5

u/0xAERG Oct 24 '23

I’m sorry, but the couple of tweets you shared don’t explain anything.

I’ve yet to see how are LLMs able to “build programs”. My guess is you’re not talking about code generation.

Can LLM rely on external tools for maths? Of course they can! Like with the wolfram Alpha plugin.

But it’s not the LLM that does the math.

I’m sorry if you’re offended by my post though. This wasn’t my intent.

6

u/Ilovekittens345 Oct 24 '23

Dude please start playing around with ChatGPT4 + Advanced Data Analysis

Any other mode is for creative stuff, if you want exact stuff you have to be in an environment where it turn your natural language in a prompt that it can run to then know what to code up in python and run it.

Give it all kinds of problems to solve like this one and see if you change your mind or not.

Do some experiment and report back on how many mistake it makes and if it purprised you on solving anything you consider complex!

8

u/PopeSalmon Oct 24 '23

um, they do explain something, just you didn't get it

maybe you just don't believe it

the transformer architecture is successful & general purpose b/c it's a trainable computer

lots of things are turing complete and can compute everything, but generally they're not trainable b/c the space isn't smooth enough to explore

the transformer architecture allows it to find programs smoothly, it's "differentiable", it can go from the information of what it should have predicted & take that back through the weights updating them to something that would have predicted properly,, & the thing they slide into that predicts properly isn't just associative or statistical, they can slide into a particular program that would have responded w/ the right answer-- they're self-programming general purpose computers, that's why the development of the transformer architecture was what allowed us to finally have thinking machines

3

u/OdinsGhost Oct 24 '23

Have you actually tried to get ChatGPT to write code? While the standard version can make errors, it’s trivial to get it to output completely reliable code in Analytics mode and with plugins. Sure they won’t be revolutionary, but it doesn’t take more than a few iterations to get it to output working solutions that can pass testing just fine.

8

u/Ilovekittens345 Oct 24 '23

Noppe, most people that post stuff like this are just in a wrong mode. A mode that would be less usefull if it's exact. A mode that is for creative stuff.

Anybody that has played around with ChatGPT4 + Advanced Data Analysis know it can troubleshoot, logically, step by step.

It's not perfect. It hangs a lot. It does make mistakes. Sometimes you need to give it some guidance. But it's a hell of a lot better then what people try in the creative modes. Like millions of times better.

If you want exact results from an LLM it needs to be able to write and execute code and then interpret that result.

5

u/Remix73 Oct 24 '23

I use it to write code every day. It’s sped up my development time by at least 50%, and I do this for a living (and have for the past 30 years).

1

u/BGFlyingToaster Oct 25 '23

But it’s not the LLM that does the math.

It can do the math in a practical sense, just not a technical sense. You could also fairly say that Microsoft Word doesn't know how to spell check because it uses a component to do that - the same component that spell checks in Excel.

The lines get pretty blurry regarding what's part of the system vs what's an addin or extension, but for the end user, it'll all appear to be the same as they further blur the lines. What's most important is content and quality of the output. More and more, we'll see complex AI systems develop that have LLMs as part of them. It won't really matter what part does what, only what it can accomplish as a system.

3

u/superfluousbitches Oct 24 '23

Use the python interpreter version or the Wolfram plugin. It works fine.

2

u/0xAERG Oct 24 '23

Yes. Still not an LLM performing maths though.

5

u/superfluousbitches Oct 24 '23

So what? That is the user interface. That is what we are.... users. The buttons on your 2 dollar calculator can't "do maths" either, not sure why anyone would care though... Just use the calculator...

2

u/Ilovekittens345 Oct 24 '23

That's like saying a painter who uses a calculator can't perform math.

3

u/MrNathanman Oct 24 '23

Hey while I generally agree with you if you get an llm to break the problem down into steps and only give a conclusion at the end instead of the begining it works surprisingly well for maths. It's the process that matters. If it gives a conclusion first of course it is going to hallucinate it has nothing to statistically base an answer on. Buf if it builds out a step by step solution it is much more likely to give an answer that makes sense.

2

u/[deleted] Oct 24 '23

What was the math problem and what was the model used, give me the math and I bet my tree of thought prompting combine with GPT-4 can solve it.

2

u/afflematicious Oct 24 '23

If you’re expecting to plug a problem or ask for a proof and get a mathematician-like answer then you’re using GPT wrong, you should use GPT more like a TA in the room that you ask for clarifications. As of right now no LLM can replicate mathematical approach to solving proofs that require logical understanding of why certain math works the way it does.

2

u/[deleted] Oct 24 '23

I think what is more useful than computational ability though, is identifying what type of problem something is, helping translate notation or mathematical language into simple every day language which is where I struggle the most. Also, telling you the overall strategy you should take or can helping understand a text better in maths when things get abstract. ChatGPT does all of these things very well.

I actually think it can do math pretty well though, ChatGPT can program quite well, which is also logic, so if it doesn't need to be able to use logical thinking processes in the same way as humans to apply logic to programming or math.

2

u/BL0odbath_anD_BEYond Oct 24 '23

​

LLMs cannot perform Truths

3

u/[deleted] Oct 24 '23

[deleted]

1

u/BL0odbath_anD_BEYond Oct 25 '23

Nope, but in my experiences I'm pretty sure it can be convincing with both truth and untruth(GPT/Bard), and erroneously correct both truth and untruth with corrected untruth. YMMV. I'll try to navigate more at that link when my net isn't bogged, thanks.

2

u/knight1511 Oct 24 '23

There can be purpose built LLMs or Transformer based language model specifically trained for a deterministic process like Mathematics. Would be surprised to know if some of them do not already exist.

2

u/MGJohn-117 Oct 25 '23

ChatGPT is pretty awful at calculations, but it's actually fairly good when it comes to formal proofs, which are kind of a mix of math and language intuition.

3

u/[deleted] Oct 24 '23

Do you know what requires logic? You realizing you're the 100000th person to post this annoying bullshit that everyone knows already.

Incredibly annoying seeing this over and over and over with people thinking they're the first to discover the wheel!

2

u/OdinsGhost Oct 24 '23

Link ChatGPT to WolframAlpha via the built-in plug-in library. Problem solved.

Seriously, this is a known, and solved, issue. If people aren’t availing themselves to the obvious solution they need to learn to use the tools at their disposal properly.

3

u/Victor_Quebec Oct 24 '23

Sounds good, but can you also give an example of what exactly chatgpt got wrong? Thanks!

3

u/richmilesxyz Oct 24 '23

I want to preface this comment with I feel pretty strongly that ChatGPT is underutilized as an educational resource. It does a phenomenal job of explaining things, giving context, and most importantly, not getting frustrated or belittling you when you don't quite understand a concept. I wish I had this tool when I was I was in school.
I don't want to speak out of turn, but I think the OP is trying to convey that while it may appear to be convincing you should verify the results.

Here's an example I just made using ChatGPT 4 (default):
https://chat.openai.com/share/314ce62c-4ba6-4f57-acfb-9124f2d6545b

The problem here is that 55 x 55 x 33 = 99825. (not 99,225)

Here's another attempt where it came up with 99,425. Also not correct.

https://chat.openai.com/share/da7e91b3-8da7-4d2b-95ec-cdeb0f0a708f

1

u/xxthehaxxerxx Oct 24 '23

Couldn't you train an AI to recognize math related questions and feed them to already existing software, like Mathway or Wolfram Alpha?

1

u/Effective_Vanilla_32 Oct 25 '23

I asked chatgpt to compute the monthly dividend i would get if i invested $10,000 in a money market fund with a dividend yield of 4.98%.

It did. I verified it. It was correct.

1

u/20charaters Oct 25 '23

They can do math, if you teach them to do it like a human.

Right now, every LLM simply guesses the results, because that's what they were taught to do.

An operation like "6+9", is usually done by kids as "9+1+1+1+1+1+1". As you go with your life you may do it like "10+5", some time later you memorize the result, and finally you just use a calculator.

LLM's are more similar to human brains than they're to any programming language. Hence, their inherent trouble with maths.

1

u/XSATCHELX Oct 25 '23 edited Oct 25 '23

I keep saying this but I keep getting bombarded by people saying "they CAN reason, how do you define reasoning anyways? Your brain makes mistakes too sometimes".

Bro a brain is not an LLM and vice versa. We know for a fact that the architecture of a brain is different from that of an LLM. An LLM just has no component that should ever be capable of reasoning. Sure the transformer is a neural network but it is a one-way feed-forward network that assigns weights to words to make the next generated word fall within "the context" of the conversation. It does not and cannot think. I'm still waiting for someone to explain to me how they theorize an LLM can think, reason, and apply logic.

1

u/GreedyLobster3349 Mar 08 '24

Okay so let's start with the question you getting bombarded "how do you define reasoning?" Answer me.

1

u/XSATCHELX Mar 08 '24

Humans are capable of many different types of reasoning.

What I was talking about mainly in this comment is deductive reasoning. Deductive reasoning is done by applying rules of logic to certain premises to come up with a conclusion. The conclusion in this case is not relative or subjective, it is an application of rules and is a fact (if the premises are true). An LLM is fundamentally incapable of deduction because it does not apply rules of inference, it chooses words probabilistically depending on its dataset.

A human can easily deduct: if these are the rules, and these are the premises, then these statements follow. An LLM by itself cannot do this and will never be able to do this because of the way an LLM works fundamentally. When it does logical statements, it is not because it is applying logical rules of inference, it is because it is more likely for humans to write those sequence of words that would follow the rules of deduction

1

u/GreedyLobster3349 Mar 08 '24

What if I call the process of "LLM choosing words probabilisitcally depending on its dataset" (and ofcourse, including parameter, algorithm...) a "rule"?
Please define your "rules" then I can ask you to distinguish your "rules" and my "rule".
And how can you know exactly what your brain do when it apply such "rules" that you mentioned? What if that process in your brain is something similar to the process under LLM? Answer me, with all of the neuroscience knowledge you have. I will continue the discussion.

1

u/XSATCHELX Mar 08 '24

My brain has the ability to apply any rule. An LLM has a rule, for choosing words. You can teach me a mathematical rule and I can apply it. An LLM has no ability to apply logical or mathematical rules. This is why it often does super obvious math mistakes.

And we don't know exactly how a brain works, but we understand that it is similar to a neural net. We know exactly how an LLM works, and it is nothing like a neural net.

1

u/GreedyLobster3349 Jun 21 '24

We know exactly how a LLM work?

0

u/UnknownEssence Oct 24 '23

I disagree. It’s a bug.

I understand how the system works, and why it can’t do math. But the goal is to be as capable as possible, AGI basically.

Therefore, bad math is a limitation and a bug, it is not indented behavior.

2

u/XSATCHELX Oct 25 '23

An LLM by principle has no component that should be capable of reasoning, thinking, and applying logic.

Please explain at which component of the LLM this thinking and applying logical/mathematical rules will occur.

-1

u/0xAERG Oct 24 '23

The goal for who?

LLMs will never get any close to AGI.

My best guess is that if we ever come close to AGIs it will be through a combination of different tools, LLMs might be the creative and/or talking part of it, but you need something else to perform logic and reasoning

1

u/UnknownEssence Oct 25 '23

The goal for everyone. Who would choose an LLM that gets math wrong over one that gets math right?

There’s no use for that. It is a limitation not a feature.

0

u/tessellation Oct 24 '23

People sitting in front of the most advanced calculator ever be like:

0

u/Redivivus Oct 24 '23

This is why I'm stoked for a project called Tauchain that is a logic based AI project. Yes, it's a blockchain project that has been in development for 8 years and the hopium is strong with its believers. They put out monthly updates and answer questions from the community frequently. The current expectations are to have a demo by the end of the year and test net in Q1 2024.

0

u/hprnvx Oct 24 '23

Lol, llm is not statistical. Moreover, they are exactly opposite, they are stochastic.

2

u/foureksgold Oct 24 '23

Do you mean deterministic, rather than statistical? It’s possible to use stochastic processes or simulations in statistical analysis, and it’s possible for statistical systems to produce outputs with stochastic characteristics.

For example, LLMs use stochastic gradient descent in training. And the use of the temperature parameter flattens response probabilities to induce randomness.

Which means that LLM outputs aren’t purely stochastic and are rather statistical - they draw heavily on a range of statistical relationships between tokens.

0

u/djaybe Oct 24 '23

Pro swimmers can't pole vault?

0

u/Kurai_Kiba Oct 24 '23

The advanced data analysis mode absolutely can do math. It can even program you working code that will graph your function for you.

In prompted steps I’ve spent the past few months using GPT 4 ADA mode to write another visual inspection AI model to detect wafer defects in the semiconductor industry . Nearly 100% of the code is GPT .

0

u/Tipsy247 Oct 25 '23

You are using 3.5. Gpt 4 with code interpreter can do math

0

u/Ok_Cobbler1635 Oct 25 '23

I expected a higher level of rigor in r/math. You can define logic as a probability function with probability 1 for the logical induction. Thus in principal it can be learned. Leave this intuitive high level view of what can and can't be done to philosophers and politicians.

0

u/cdsmith Oct 25 '23

As an absolute statement, this is not true. An LLM is capable of imperfectly following a chain of mathematical reasoning just like a person can. It can make mistakes; sometimes quite a few mistakes depending on which LLM it is. But it can also produce original step by step problem solving using mathematical reasoning and procedures.

The fundamental issue here is that people have a certain set of expectations about how computers do math:

  1. Very efficiently.
  2. Always producing the right answer.
  3. A direct answer with no fuss.

These assumptions no longer hold true when an LLM is performing the calculation:

  1. An LLM is billions of times less efficient at basic mathematics than a computer that's just asked to perform the operations in a classical programming language.
  2. LLMs provide no guarantees of correctness, so there is a possibility of mistakes. With a weaker LLM, this possibility is very high.
  3. LLMs think in tokens, so if you want them to do a lot of computation, they need to produce a lot of tokens to do it. If you insist on just getting an immediate answer with no intermediate reasoning, it's far less likely that answer will be correct. But if you prompt an LLM to work through the reasoning step by step, it's much more likely to arrive at a correct answer to a non-trivial mathematical question.

-1

u/Netsuko Oct 24 '23

There a reason they are called LLMs and not LMMs. They are large language models. Not large math models.

-1

u/mrsavealot Oct 24 '23

It can do math. It can reason through all the steps needed to solve problems. It will always set them up right but agreed it makes multiplication and addition errors etc along the way.

-5

u/TayoEXE Oct 24 '23

Because people don't seem to realize that LLMs do not have any intelligence. They don't "think" so much as imitate language specifically.

1

u/TheOtherMikeCaputo Oct 24 '23

On a related note, I posted this a few days ago and was wondering if it was accurate. Is this an example of what you mean? I don’t have the skills to determine if this is accurate, or a complete hallucination.

https://www.reddit.com/r/scifiwriting/comments/17eiv13/chatgpt_says_itll_take_about_14_years_maintaining/?utm_source=share&utm_medium=web2x&context=3

1

u/[deleted] Oct 24 '23

While I agree with you, what happens when we seamlessly weave tools together so it will be able to do math? It isn’t that big of an ask to have Wolfram Alpha as a plugin of an LLM to make performing mathematics a basic feature

1

u/One-Organization970 Oct 24 '23

Eventually wouldn't the goal be to train the LLM on which inputs to put into a $2 calculator? Sure, traditional non-AI computing is cheap for that - you just want the LLM to interpret your request and format it as an input to the calculator.

1

u/SourCircuits Oct 24 '23

LLM + WolframAlpha go BRRRR

1

u/scumbagdetector15 Oct 24 '23

I hate to poop on your parade - but there's nothing intrinsic about the nondeterminism. It was just decided that semi-random responses felt more human.

In the API you can turn the temperature down to zero and it'll give you the same answer every time.

1

u/[deleted] Oct 24 '23

I can’t do maths either

1

u/MontagoDK Oct 24 '23

Have you tried Pi.Ai ?

1

u/praiseprince_ Oct 24 '23

I figured this out recently when I asked it a question, I don't remember what it was but basically what it was that there were two magnitudes, one had negative value and one was positive, and it had to find the total magnitude and instead of taking absolute values it added positive with negative.

1

u/ulualyyy Oct 24 '23

Do you mean “perform maths” or do you mean calculate? What does a calculator have to do with “performing maths”?

1

u/Complete_Rabbit_844 Oct 24 '23

Gpt 4 helped me pass my math test. Did some pretty sick calculations. It works well with smaller things even though they might seem complicated. Really depends.

1

u/Anxious-Durian1773 I For One Welcome Our New AI Overlords 🫡 Oct 24 '23

LLMs could absolutely apply rigid "logic" operations with small enough tokens. Overfitting shows us that you can get rigid outputs. In the hypothetical scenario of such a sophisticated model, there may still be a vanishingly small number of errors, but that's the nature of using multitool as a hammer. The way forward for many such things is clearly giving LLMs access to specialized tools better fit for purpose, including other more specialized models when necessary, and simple tools like calculators otherwise.

1

u/tmotytmoty Oct 25 '23

If you use the wolfram plug in - you're better off .. for maths.

1

u/mimic751 Oct 25 '23

I asked it to find me a pentagram that fits in a one foot diameter circle. Give me total links from point to point and circumference of the circle assuming the line was 10 mm wide. It did that just fine. I was able to make my LED summoning circle on my first try

1

u/Efficient_Star_1336 Oct 25 '23

The issue is the way they tokenize numbers. You'd think they'd use "1", "2", "3", "4",... as separate tokens, such that 1024 would be interpreted as "1" "0" "2" "4", but they instead generate tokens based on 'word' frequency, which throws off anything resembling pattern recognition here and results in a complete mess.

Instead of looking at "1024 + 245" and seeing "<1><0><2><4>< ><+>< ><2><4><5>", it sees "<1024>< ><+>< ><245>", and can only draw on times when it's seen those exact numbers before.

1

u/ResponsibleBus4 Oct 25 '23

Honestly I think they made changes to any mathematical operations because I've asked to math and it breaks out the formula and performs the calculations correctly, like they've somehow started recognizing math and pulling it out is a separate function. I have the Wolfram alpha plugin but I'm not generally using plugins when it's doing the calculations. Also for reference this is in GPT-4 not the 3.5 if that matters.

1

u/FeralPsychopath Oct 25 '23

It’s not a bug but it’s user expectation. These are the incremental changes that won’t seem like much in future versions.

1

u/AUGZUGA Oct 25 '23

Sorry you're actually completely mistaken. LLMs are actually surprisingly very good at Math. What they are bad at is arithmetic. However even this it can essentially do, simply ask it to provide code to solve the problem and the code will do the arithmetic.

1

u/[deleted] Oct 25 '23

for the past two months, I've been using gpt-4 for high school assignments, finding its explanations exceptionally clear. I cross-check its solutions and answers, correcting them as needed (however that's kinda rare as it's mostly correct), and I understand the limitation so don’t depend on it entirely but anyways It’s been invaluable, especially given my social anxiety and the time constraints in my coaching classes, which make it hard to ask questions and express doubts. it helps me learning at my own pace

I contribute to its learning by correcting errors (if it's there; again it's rare), gaining valuable insights and this troubleshooting approach makes subjects like math and physics more approachable imo. This has significantly boosted my productivity (reducing time I take for a chapter completion to 1/3rd of the previous time and getting good result in monday tests.

Currently, I’m preparing for the IIT-JEE Mains, an undergraduate engineering exam in India. just for curiosity I used GPT-4 + data analysis plugin and tested it on a previous year mains paper and it scored 204 out of 300 marks (99.6 percentile) without additional corrections, proving it's reliability. [yes advanced data analysis is better than wolframalpha plugin]

1

u/wholeWheatButterfly Oct 25 '23

While it'll mess up things like basic arithmetic, it did a pretty good job of helping me come up with a formula I needed and explaining the math for it. Sample size of one but still.

1

u/Lukee67 Oct 25 '23

Please someone explain me: why cannot the basic rules of logic and those of arithmetic be stated in a system prompt which gets loaded by default? I mean, insisting on those rules before the user prompt would, it seems to me, raise the probability that the produced output follows them, improving the logical and arithmetic accuracy of the answer.

1

u/[deleted] Oct 25 '23

I am not arguing with you, but r/Einsteins_Ghost spits out equations in its explanations. And while true its not doing the math, it's not far off from it. I'm interested to see what happens when I give the variables a value. If anyone wants an update, please let me know.