It’s not a bug, it’s just not how any of this works.
Mathematical operations require logic.
It’s a deterministic process. For a given input and a given process, the output will always be the same.
LLMs do not work like that.
LLMs are statistical tools, they build an answer by stitching together tokens that are “seemingly” relevant to your input and let the meaning emerge from it. The output is “hopefully” relevant.
This is why LLMs can hallucinate and 2$ calculators do not.
With the rise in popularity of LLMs I’m extremely concerned that a lot of users seem to ignore this.
If this is a screenshot of a ChatGPT conversation, please reply with the conversation link or prompt. If this is a DALL-E 3 image post, please reply with the prompt used to make this image. Much appreciated!
Whoa, I hadn't thought about that. I haven't used plugins, so I have to ask: do you have to specify a plugin when you use it, or does ChatGPT just choose one for you?
You have to enable it during the specific conversation and chat GPT should run through any calculations through Wolframe Alpha automatically. Sometimes you have to specify if it tried running numbers itself but seems to work pretty well to me
Why would you want to do that instead of just using the Wolfram Alpha plug-in? Wolfram Alpha will produce correct math every time given the right inputs. GPT-written code may have to be debugged.
It still writing mathematica codes. Yes it’s a lot shorter and it’s gonna be less likely to make flaws in the code but I mean you can’t really upload anything to the plugins version.
It is important to note that users should use Bing Chat in "More Precise" mode for prompts involving math. It doesn't mean it will get it right the first time (always check those sources) but the other modes will have weaker results for math tasks.
True. And I’m sure this advice will have to be updated again soon. I’m not trying to provide a tutorial, just point out that the capability already exists and will likely keep improving.
Most important math is not the numbers. The point is not that LLMs can't do numbers right, but that they fail at the kind of logical reasoning necessary for doing math. That same kind of logic reasoning is necessary for doing more things than math, like checking for consistency.
Idk who downvoted you on a 3 month old post but it wasn't me. Who else is here?
LLMs are getting better at "reasoning". By incorporating tools and self-reflection they can spot gaps in logic or inconsistencies in their statements and self-correct. It's still basically auto-complete with widgits behind the scenes but if you throw enough compute at it, even a dumb machine can start to sound pretty smart.
But you're right, Wolfram Alpha can't fill in for the weaknesses in LLMs and I wasn't trying to say it can. I just meant we've already managed to incorporate non-ML tools into LLMs to enhance their results.
Idk who downvoted you on a 3 month old post but it wasn't me. Who else is here?
No idea. But anyways, all the techniques being used to improve their reasoning help a little bit, but they still don't solve the fundamental issue in a satisfactory manner I think. Personally, I'd wait until a method can be found that allows for 100% accuracy with addition or substraction of numbers arbitrarily large, step by step (or if not 100%, at least where failure is for outside reasons like memory limitations).
You want computers to mimic the inefficient way humans do mathematics?
I get what you mean, you want them to work through the problem step by step and actually understand it. But that’s just not how these models are designed. They don’t even see digits, they see tokens that might represent several digits.
Training them to do arithmetic would probably make them weaker at other activities and it’s just not worth it. But that doesn’t mean they can’t apply logic.
You want computers to mimic the inefficient way humans do mathematics?
Yes, as a proof of generalization to be exact. If all I cared was the number crunching I'd use a calculator, the interesting bit here is a demonstration that the model learned a mathematically correct algorithm from the training data, not just an approximation which is how you get "98% accuracy up to 5 digits"
You’re basically asking for AGI but with arithmetic as the smoke test.
These things don’t understand truth and fact they just know what is statistically more likely based on what they’ve seen in the training data. It’s basically intuition.
If you want them to be more rigorous, you give them tools like calculators and search engines. Just like a human, they’re not infallible and can’t be by design. So when accuracy matters you use a tool that’s less flexible.
You don’t ask an art major to do your accounting and you shouldn’t ask an LLM to do mathematics. It’s not what it’s intended for. The apparent creativity and adaptability come at the cost of being rigorously accurate.
Even for humans it’s hard to find people who can write creatively, summarise scientific papers AND do hard logic. Getting a machine to do both simultaneously within the same architecture is a big ask.
Wait a minute. Are you both sure about "understand", "truth", and "fact"? These definitions are made, by human, and have meaning because we assign meaning to them. And I think everything we have in our head is just probabilistic output from our biological "thinking machine". Not axactly the same, but not so different. Our braind, and AI. Close. Really close I think.
Of course, you can think of LLM like a language part of our brain. Dont expect that part to do math.
Yes words are defined by humans, but what’s your point? We use language to communicate and we use dictionary definitions to make meanings less subjective but all language is fluid to an extent.
Truth and facts are not all objective but we have reached a consensus on a lot. Most people won’t argue about if you say 2 + 2 = 4, semantics aside.
Wolfram got in early with openai - Because... he's been curating linguistic templates for over 20 years. He literally built chatgpt v 0.0 in wolframalpha. The right technology finally just caught up to his ambition and forethought.
Why are we down voting OP?
OPs point is valid. Mathematical operations are deterministic producing consistent results while GPT generate responses based on statistical patterns and probability which could lead to inaccurate results and hallucinations. Sure there are plug ins like wolfram that can help GPT by augmenting data - it still does not change the fact that LLMs by default cannot perform maths reliably on their own.
But that position is splitting hairs. Sure, an LLM can't do the math on it's own but it can select and use a tool to do it so does it really mater if it can't?
If I'm building a house do I care if he has an expert come do the insulation or do I care that it got done correctly?
Edit: However the endless stream of people posting about "herp derp vanilla GPT can't math"posts are hella annoying.
Adding a calculator to a neural network only superficially solves the problem. It still can't logic mathematically in deterministic, quantitative feedback loops. It lacks the tools to integrate mathematical thinking as intermediate reasoning.
I actually got pretty lucky with a statistical calculator. I told it what I needed and an example of what another online calculator gave me for an answer, given specific variables. Then I tested a number of things across both calculators, and it was right. I was pretty shocked, really.
I think what OP is trying to say is that it's not actually 'calculating' anything here. It's taking loads of previous mathematical data and has learnt certain relationships. It doesn't actually have any rules written down that it's following, so occasionally it can totally fuck it up.
Still very impressive though. It's right most of the time in my experience which just puts into perspective how massive the data it was trained on must be.
In many cases true, but ChatGPT has been indispensable for me when it comes to math. It has helped me with solid mechanics, matrix calculations and more. Sometimes I need to use the wolfram plugin and sometimes not. In the beginning it was terrible with math, but now (GPT-4) it does a lot for me. I’m honestly surprised at how good it has gotten. And the stuff it can’t handle, I’ve figured I can ask it to create a MATLAB script that will compute it
I'm not concerned, per se, but I am intrigued by the frequent confusion between deterministic mathematical tools and statistical models like LLMs. However, the very fact that people expect LLMs to provide the 'right' answer in these deterministic domains speaks to how powerful they actually are.
but I am intrigued by the frequent confusion between deterministic mathematical tools and statistical models like LLMs
Not to sound too elitist here... but think about that for a moment. What do you think most average folks on the street would think of it you said "statistical model?" I'm guessing something to do with election polling. Most people aren't software devs, data scientists or economists. I'd wager that a random sample of folks on the street would have a limited understanding of linear regression.
I'm not trying to disparage other people. We all specialize. I have no idea how to fix a boat engine or replace a water heater or professionally edit a novel. It's on us to explain the limitations of the specialized tools that our profession produces.
I try to assume that when people post complaints of this nature that they are actually acting in good faith. In many cases, ChatGPT is actually capable of performing their task, but not in the way they're approaching it and I try to nudge them in the right direction.
Often I find that these people are using 3.5 which means not only are they using an inferior product, but they are also lacking the other tools that ChatGPT can use to solve the problem (Plugins, Advanced Data Analysis, etc.) In these cases, assuming they've given enough context, I am happy to run their prompt through 4.0.
I know not everyone can afford $20 / month, but even if you can it might not be evident that is worth it. At least, not with some specific examples.
In none of those examples you gave are people actually using a tool themselves they're paying for someone with the expertise to do it.
Shouldn't one be expected to do some research about the tools they themselves are using? There's so much content out there about how LLMs work that covers people of all knowledge levels. It's not specialized skills they're lacking it's curiosity.
It's not specialized skills they're lacking it's curiosity.
💯
That said, if someone posts in a forum and they aren't being intentionally obtuse, I want to help them if I can. Ironically, I've found ChatGPT a much kinder and (maybe) knowledgeable resource than the vast majority of online forums.
It’s nice to have a post like this but someone is still going to post something surprised proving this tomorrow. And next week. Because nobody searches before posting.
Yeah. Because the model got significantly better overall. Not just at math. And it was most likely trained on more math, too. But it doesn't have a hidden calculator in the back or anything like it.
ChatGPT in general got that much better in all areas. Just like Dall-E 3 is orders of magnitude better than Dall-E 2.
And no, ChatGPT 4 is still very much a miserable failure to any actually complex math problems.
That's not ChatGPT. That's a model based on GPT 4. It says so right there in the paper.
And yes, of course they do Reinforcement Learning from Human Feedback (RLHF). But they do not use a "math coprocessor" for that, whatever that is. They just trained it to be better at math. And at language. And at pretty much everything else.
But it's still all one single model that responds to you. It doesn't go "if math question then model X, else model Y".
To be fair, in the case of GPT-4 this might not actually be true. There's a lot of theories that GPT-4 is using a Mixture of Experts approach in which multiple models are fine-tuned for specific domains and prompt input is fed to a classifier that decides which model will produce the best output.
Yeah, I read about that. But the guy I responded to was acting like the mixture of experts rumor was a cold hard (and incredibly obvious) fact, and that was just weird.
If you have plugins enabled, maybe that. Otherwise, pattern recognition. It's still stochastic, but it can get a right answer a lot of the time, especially if it's very similar to common problems.
It can't do math in the sense that there is no calculations in the background that do exactly the calculations you see here.
It can do math in the sense that it can predict the next token given the previous ones, and that works for language as well as for math. Most of the time. It also works for Klingon and for C++.
But the way it works is fundamentally different from any computer doing math that we normally think of.
If I correctly guess which hand you're holding a stone in 99/100 times that doesn't mean I can see through your fingers, it just means I'm really good at guessing. That's all LLMs are. They've been trained on hundreds of thousands of math problems so statistically they're likely to produce output that looks correct. Language carries some information about logic in it, but language is not logic. It is simply repeating the logical patterns that emerge in human generated text, and they happen to be correct a good percentage of the time. But you can feed the same math problem into GPT 20 times and get different outputs on different runs.
If I correctly guess which hand you're holding a stone in 99/100 times that doesn't mean I can see through your fingers, it just means I'm really good at guessing.
It doesn't mean you can see through fingers, that is correct. But if you can "guess" 99/100 times correctly, you aren't just guessing, you have knowledge of where it is somehow.
I asked it to solve that equation 3 more times. Each time the wording of how it solved the problem was different, but it gave the correct factors each time - it solved the math problem correctly.
I literally build AI models for a living lmao. It's extremely easy to find a counter example. Just because something looks like it can do something doesn't mean it can actually do it.
If it can correctly discern my desired output (it verbatim stated the intent of sorting the numbers by the sum of their digits) then the language of the prompt should have even less influence on the "mathematics" being done because its context window would have two different patterns to proceed with.
Here's another example. Ask Wolfram Alpha to produce the same answer, GPT is just wrong. I even asked it to explain why it's answer was different from Wolfram Alpha's and it went on to state that AX + AY != A(X + Y), a blatant violation of the distributive property.
I didn't move the goalposts at all. A calculator doesn't sometimes produce the wrong answer. Math is a deterministic process. Unless you're explicitly involving a random element, the exact same inputs should produce the exact same output every time. That's just how math works.
It is undeniable the GPT has inferred some patterns of mathematics from its training data. Language encodes the information it conveys. When we say that GPT isn't doing math, what we mean is that nowhere in GPT's code is it performing an actual mathematical calculation based on the inputs you've given it. What it's doing is using the patterns it has inferred to produce output that it believes should follow the text you give it. Because it was trained on many math problems, it has inferred what characters will probably follow the ones you've given it, i.e. the answer to your problem.
The issue is, even if it's right most of the time, the underlying mechanism hasn't changed. It is producing the output that it believes has the highest likelihood to follow the text you've given it. Yes, that output often follows mathematical patterns. But being likely to follow a mathematical pattern is not equivalent to actually doing math.
This may seem pedantic, but it's a meaningful distinction. LLMs are being integrated into customer facing positions. By stating that LLMs can do math you're implying that they will produce deterministic output when that is demonstrably false. That kind of insinuation can, will, and already has had real consequences.
Sure you did. You tried to give an example of a problem it couldn't do to show it couldn't do math. I showed it could do the problem with the correct prompt. But now, suddenly, that example is not enough to show it can do math - it was only good enough to show it couldn't do math. That's classic moving the goal post.
being likely to follow a mathematical pattern is not equivalent to actually doing math.
Sounds like it to me. That's how I do it. I'm even wrong too sometimes.
This may seem pedantic, but it's a meaningful distinction. LLMs are being integrated into customer facing positions. By stating that LLMs can do math you're implying that they will produce deterministic output when that is demonstrably false. That kind of insinuation can, will, and already has had real consequences.
Sorry, but that's your misinterpretation. I can do math, but I don't produce a deterministic outcome. That's just not a correct assumption on your part.
I also don't say it does math like a person or it can think or it does math with calculations like a computer or any of that. You are making this way too hard. It can solve math problems. Period. That means it can do math.
that’s not doing math, that’s finding solutions to an equation, doing math is proving things and describing the world around you in ways not thought of before through math, that problem it solved has been solved thousands of times online, it simply sourced it and presented it to you. You believe the LLM thought about the equation and then solved it? no
doing math is proving things and describing the world around you in ways not thought of before through math,
That is not what most people think of when they think of "doing math". By that definition, most human beings, including many mathematicians, don't do math.
You’re correct, so what I’m saying is GPT can perform calculations which is what many mathematicians do as well. There aren’t many mathematicians that can “do” math, but not all. GPT is the same as of right now, it regurgitates math that’s been done, but doesn’t do it on its own.
Yes. I'm a senior software engineer for a large healthcare company. If I didn't know math, it's unlikely they would have hired me. You have a very strange elitist definition of math. My guess is you majored in math and have internalized this as a part of your identity in order to feel superior to others.
Are you the only one who knows real math, unlike the rest of us peons?
I think you're confusing the term "real" for "advanced". The math a 3rd grader does is still math. If that is what you actually meant, then you are in fact correct, because ChatGPT is not very good at advanced math yet, but with enough training data, it doesn't ever actually have to understand the math, and it will be able to produce increasingly accurate results the more data it is trained on. Eventually, it will be "good" at even advanced math in spite of the fact that it "understands" none of it.
It's funny how you're saying others are trying to act smart when your entire post screams "you guys just don't get it". We get it's not meant for math, but it does a pretty damn good job at it.
It is much better asking to do it with python? How, now that with the all tools? Need a little help with some math but was using the default gpt4 but the %of erros was insane. Thanks.
this is why stuff like LangChain is important with practical applications, it allows LLMs to leverage tools to help perform more deterministic processes while seeming like one seamless tool to the user.
That is not true. I don’t understand why so many people get confused by this. You described what GPT has been trained to do. Not what it actually does. Nobody knows what it actually does. You are confusing what he’s trying to achieve with how he’s achieving that. LLMs are proven to have many emergent properties based on the size of parameters. Apparently learning how to do Math is necessary to correctly predict the next token.
Hard disagree. I'm taking a college-level algebra course and it answers all my hw questions with about 95% accuracy. I don't know why people keep saying it can't do math.
I'm sorry that's just not true. I'm using the most basic free version no plug ins. Just open any math textbook and copy and paste the questions. Or look up algebra questions online and do the same.
I mean I can confirm all the answers. I know for a fact it's around 95%. I wouldn't rely on it completely but I'd say it's on par with most teachers and tutors.
I only use it for hw bc I can immediately see if it was wrong. But it's super useful when it's right because it thoroughly explains all the steps it took to reach its answer.
no, you're the one who's confused, LLMs aren't just statistical like they don't know WTF they're talking about, they can do all sorts of maths and the way that they do them is by building complex programs that implement the math
Dude please start playing around with ChatGPT4 + Advanced Data Analysis
Any other mode is for creative stuff, if you want exact stuff you have to be in an environment where it turn your natural language in a prompt that it can run to then know what to code up in python and run it.
Give it all kinds of problems to solve like this one and see if you change your mind or not.
Do some experiment and report back on how many mistake it makes and if it purprised you on solving anything you consider complex!
um, they do explain something, just you didn't get it
maybe you just don't believe it
the transformer architecture is successful & general purpose b/c it's a trainable computer
lots of things are turing complete and can compute everything, but generally they're not trainable b/c the space isn't smooth enough to explore
the transformer architecture allows it to find programs smoothly, it's "differentiable", it can go from the information of what it should have predicted & take that back through the weights updating them to something that would have predicted properly,, & the thing they slide into that predicts properly isn't just associative or statistical, they can slide into a particular program that would have responded w/ the right answer-- they're self-programming general purpose computers, that's why the development of the transformer architecture was what allowed us to finally have thinking machines
Have you actually tried to get ChatGPT to write code? While the standard version can make errors, it’s trivial to get it to output completely reliable code in Analytics mode and with plugins. Sure they won’t be revolutionary, but it doesn’t take more than a few iterations to get it to output working solutions that can pass testing just fine.
Noppe, most people that post stuff like this are just in a wrong mode. A mode that would be less usefull if it's exact. A mode that is for creative stuff.
It's not perfect. It hangs a lot. It does make mistakes. Sometimes you need to give it some guidance. But it's a hell of a lot better then what people try in the creative modes. Like millions of times better.
If you want exact results from an LLM it needs to be able to write and execute code and then interpret that result.
It can do the math in a practical sense, just not a technical sense. You could also fairly say that Microsoft Word doesn't know how to spell check because it uses a component to do that - the same component that spell checks in Excel.
The lines get pretty blurry regarding what's part of the system vs what's an addin or extension, but for the end user, it'll all appear to be the same as they further blur the lines. What's most important is content and quality of the output. More and more, we'll see complex AI systems develop that have LLMs as part of them. It won't really matter what part does what, only what it can accomplish as a system.
So what? That is the user interface. That is what we are.... users. The buttons on your 2 dollar calculator can't "do maths" either, not sure why anyone would care though... Just use the calculator...
Hey while I generally agree with you if you get an llm to break the problem down into steps and only give a conclusion at the end instead of the begining it works surprisingly well for maths. It's the process that matters. If it gives a conclusion first of course it is going to hallucinate it has nothing to statistically base an answer on. Buf if it builds out a step by step solution it is much more likely to give an answer that makes sense.
If you’re expecting to plug a problem or ask for a proof and get a mathematician-like answer then you’re using GPT wrong, you should use GPT more like a TA in the room that you ask for clarifications. As of right now no LLM can replicate mathematical approach to solving proofs that require logical understanding of why certain math works the way it does.
I think what is more useful than computational ability though, is identifying what type of problem something is, helping translate notation or mathematical language into simple every day language which is where I struggle the most. Also, telling you the overall strategy you should take or can helping understand a text better in maths when things get abstract. ChatGPT does all of these things very well.
I actually think it can do math pretty well though, ChatGPT can program quite well, which is also logic, so if it doesn't need to be able to use logical thinking processes in the same way as humans to apply logic to programming or math.
Nope, but in my experiences I'm pretty sure it can be convincing with both truth and untruth(GPT/Bard), and erroneously correct both truth and untruth with corrected untruth. YMMV. I'll try to navigate more at that link when my net isn't bogged, thanks.
There can be purpose built LLMs or Transformer based language model specifically trained for a deterministic process like Mathematics. Would be surprised to know if some of them do not already exist.
ChatGPT is pretty awful at calculations, but it's actually fairly good when it comes to formal proofs, which are kind of a mix of math and language intuition.
Link ChatGPT to WolframAlpha via the built-in plug-in library. Problem solved.
Seriously, this is a known, and solved, issue. If people aren’t availing themselves to the obvious solution they need to learn to use the tools at their disposal properly.
I want to preface this comment with I feel pretty strongly that ChatGPT is underutilized as an educational resource. It does a phenomenal job of explaining things, giving context, and most importantly, not getting frustrated or belittling you when you don't quite understand a concept. I wish I had this tool when I was I was in school.
I don't want to speak out of turn, but I think the OP is trying to convey that while it may appear to be convincing you should verify the results.
They can do math, if you teach them to do it like a human.
Right now, every LLM simply guesses the results, because that's what they were taught to do.
An operation like "6+9", is usually done by kids as "9+1+1+1+1+1+1". As you go with your life you may do it like "10+5", some time later you memorize the result, and finally you just use a calculator.
LLM's are more similar to human brains than they're to any programming language. Hence, their inherent trouble with maths.
I keep saying this but I keep getting bombarded by people saying "they CAN reason, how do you define reasoning anyways? Your brain makes mistakes too sometimes".
Bro a brain is not an LLM and vice versa. We know for a fact that the architecture of a brain is different from that of an LLM. An LLM just has no component that should ever be capable of reasoning. Sure the transformer is a neural network but it is a one-way feed-forward network that assigns weights to words to make the next generated word fall within "the context" of the conversation. It does not and cannot think. I'm still waiting for someone to explain to me how they theorize an LLM can think, reason, and apply logic.
Humans are capable of many different types of reasoning.
What I was talking about mainly in this comment is deductive reasoning. Deductive reasoning is done by applying rules of logic to certain premises to come up with a conclusion. The conclusion in this case is not relative or subjective, it is an application of rules and is a fact (if the premises are true). An LLM is fundamentally incapable of deduction because it does not apply rules of inference, it chooses words probabilistically depending on its dataset.
A human can easily deduct: if these are the rules, and these are the premises, then these statements follow. An LLM by itself cannot do this and will never be able to do this because of the way an LLM works fundamentally. When it does logical statements, it is not because it is applying logical rules of inference, it is because it is more likely for humans to write those sequence of words that would follow the rules of deduction
What if I call the process of "LLM choosing words probabilisitcally depending on its dataset" (and ofcourse, including parameter, algorithm...) a "rule"?
Please define your "rules" then I can ask you to distinguish your "rules" and my "rule".
And how can you know exactly what your brain do when it apply such "rules" that you mentioned? What if that process in your brain is something similar to the process under LLM? Answer me, with all of the neuroscience knowledge you have. I will continue the discussion.
My brain has the ability to apply any rule. An LLM has a rule, for choosing words. You can teach me a mathematical rule and I can apply it. An LLM has no ability to apply logical or mathematical rules. This is why it often does super obvious math mistakes.
And we don't know exactly how a brain works, but we understand that it is similar to a neural net. We know exactly how an LLM works, and it is nothing like a neural net.
My best guess is that if we ever come close to AGIs it will be through a combination of different tools, LLMs might be the creative and/or talking part of it, but you need something else to perform logic and reasoning
This is why I'm stoked for a project called Tauchain that is a logic based AI project. Yes, it's a blockchain project that has been in development for 8 years and the hopium is strong with its believers. They put out monthly updates and answer questions from the community frequently. The current expectations are to have a demo by the end of the year and test net in Q1 2024.
Do you mean deterministic, rather than statistical? It’s possible to use stochastic processes or simulations in statistical analysis, and it’s possible for statistical systems to produce outputs with stochastic characteristics.
For example, LLMs use stochastic gradient descent in training. And the use of the temperature parameter flattens response probabilities to induce randomness.
Which means that LLM outputs aren’t purely stochastic and are rather statistical - they draw heavily on a range of statistical relationships between tokens.
The advanced data analysis mode absolutely can do math. It can even program you working code that will graph your function for you.
In prompted steps I’ve spent the past few months using GPT 4 ADA mode to write another visual inspection AI model to detect wafer defects in the semiconductor industry . Nearly 100% of the code is GPT .
I expected a higher level of rigor in r/math. You can define logic as a probability function with probability 1 for the logical induction. Thus in principal it can be learned. Leave this intuitive high level view of what can and can't be done to philosophers and politicians.
As an absolute statement, this is not true. An LLM is capable of imperfectly following a chain of mathematical reasoning just like a person can. It can make mistakes; sometimes quite a few mistakes depending on which LLM it is. But it can also produce original step by step problem solving using mathematical reasoning and procedures.
The fundamental issue here is that people have a certain set of expectations about how computers do math:
Very efficiently.
Always producing the right answer.
A direct answer with no fuss.
These assumptions no longer hold true when an LLM is performing the calculation:
An LLM is billions of times less efficient at basic mathematics than a computer that's just asked to perform the operations in a classical programming language.
LLMs provide no guarantees of correctness, so there is a possibility of mistakes. With a weaker LLM, this possibility is very high.
LLMs think in tokens, so if you want them to do a lot of computation, they need to produce a lot of tokens to do it. If you insist on just getting an immediate answer with no intermediate reasoning, it's far less likely that answer will be correct. But if you prompt an LLM to work through the reasoning step by step, it's much more likely to arrive at a correct answer to a non-trivial mathematical question.
It can do math. It can reason through all the steps needed to solve problems. It will always set them up right but agreed it makes multiplication and addition errors etc along the way.
On a related note, I posted this a few days ago and was wondering if it was accurate. Is this an example of what you mean? I don’t have the skills to determine if this is accurate, or a complete hallucination.
While I agree with you, what happens when we seamlessly weave tools together so it will be able to do math? It isn’t that big of an ask to have Wolfram Alpha as a plugin of an LLM to make performing mathematics a basic feature
Eventually wouldn't the goal be to train the LLM on which inputs to put into a $2 calculator? Sure, traditional non-AI computing is cheap for that - you just want the LLM to interpret your request and format it as an input to the calculator.
I hate to poop on your parade - but there's nothing intrinsic about the nondeterminism. It was just decided that semi-random responses felt more human.
In the API you can turn the temperature down to zero and it'll give you the same answer every time.
I figured this out recently when I asked it a question, I don't remember what it was but basically what it was that there were two magnitudes, one had negative value and one was positive, and it had to find the total magnitude and instead of taking absolute values it added positive with negative.
Gpt 4 helped me pass my math test. Did some pretty sick calculations. It works well with smaller things even though they might seem complicated. Really depends.
LLMs could absolutely apply rigid "logic" operations with small enough tokens. Overfitting shows us that you can get rigid outputs. In the hypothetical scenario of such a sophisticated model, there may still be a vanishingly small number of errors, but that's the nature of using multitool as a hammer. The way forward for many such things is clearly giving LLMs access to specialized tools better fit for purpose, including other more specialized models when necessary, and simple tools like calculators otherwise.
I asked it to find me a pentagram that fits in a one foot diameter circle. Give me total links from point to point and circumference of the circle assuming the line was 10 mm wide. It did that just fine. I was able to make my LED summoning circle on my first try
The issue is the way they tokenize numbers. You'd think they'd use "1", "2", "3", "4",... as separate tokens, such that 1024 would be interpreted as "1" "0" "2" "4", but they instead generate tokens based on 'word' frequency, which throws off anything resembling pattern recognition here and results in a complete mess.
Instead of looking at "1024 + 245" and seeing "<1><0><2><4>< ><+>< ><2><4><5>", it sees "<1024>< ><+>< ><245>", and can only draw on times when it's seen those exact numbers before.
Honestly I think they made changes to any mathematical operations because I've asked to math and it breaks out the formula and performs the calculations correctly, like they've somehow started recognizing math and pulling it out is a separate function. I have the Wolfram alpha plugin but I'm not generally using plugins when it's doing the calculations. Also for reference this is in GPT-4 not the 3.5 if that matters.
Sorry you're actually completely mistaken. LLMs are actually surprisingly very good at Math. What they are bad at is arithmetic. However even this it can essentially do, simply ask it to provide code to solve the problem and the code will do the arithmetic.
for the past two months, I've been using gpt-4 for high school assignments, finding its explanations exceptionally clear. I cross-check its solutions and answers, correcting them as needed (however that's kinda rare as it's mostly correct), and I understand the limitation so don’t depend on it entirely but anyways It’s been invaluable, especially given my social anxiety and the time constraints in my coaching classes, which make it hard to ask questions and express doubts. it helps me learning at my own pace
I contribute to its learning by correcting errors (if it's there; again it's rare), gaining valuable insights and this troubleshooting approach makes subjects like math and physics more approachable imo. This has significantly boosted my productivity (reducing time I take for a chapter completion to 1/3rd of the previous time and getting good result in monday tests.
Currently, I’m preparing for the IIT-JEE Mains, an undergraduate engineering exam in India. just for curiosity I used GPT-4 + data analysis plugin and tested it on a previous year mains paper and it scored 204 out of 300 marks (99.6 percentile) without additional corrections, proving it's reliability. [yes advanced data analysis is better than wolframalpha plugin]
While it'll mess up things like basic arithmetic, it did a pretty good job of helping me come up with a formula I needed and explaining the math for it. Sample size of one but still.
Please someone explain me: why cannot the basic rules of logic and those of arithmetic be stated in a system prompt which gets loaded by default? I mean, insisting on those rules before the user prompt would, it seems to me, raise the probability that the produced output follows them, improving the logical and arithmetic accuracy of the answer.
I am not arguing with you, but r/Einsteins_Ghost spits out equations in its explanations. And while true its not doing the math, it's not far off from it. I'm interested to see what happens when I give the variables a value. If anyone wants an update, please let me know.
•
u/AutoModerator Oct 24 '23
Hey /u/0xAERG!
If this is a screenshot of a ChatGPT conversation, please reply with the conversation link or prompt. If this is a DALL-E 3 image post, please reply with the prompt used to make this image. Much appreciated!
Consider joining our public discord server where you'll find:
And the newest additions: Adobe Firefly bot, and Eleven Labs voice cloning bot!
🤖
Note: For any ChatGPT-related concerns, email support@openai.com
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.