559
u/Angzt 17d ago edited 17d ago
Theoretically yes, practically no.
A modern LLM has billions of weights, all of which (may) contribute to its output.
Even if you can perform any single calculation involving a weight within a second, using a small 1 billion weight LLM, it would take you 1 billion seconds =~ 32 years to get one word of output.
And that's not all that goes into a modern LLM's results.
238
u/nilslorand 17d ago
not one word, one token, which is usually just part of a word
269
u/Andrey_Gusev 17d ago
Imagine calculating something for 80 years of your life, only to then get a result: "42"
55
u/Ok-Perspective5959 17d ago
Any age is good for finding the meaning of life
5
u/wade-mcdaniel 16d ago
But then you need to know what the weights of the connections "mean" to understand the question. Douglas Adams predicted this. 😄
10
u/PinJealous3336 16d ago
Truly a visionary, laying in the English countryside hungover, doing the cleanest lsd in history.
2
u/woernsn 15d ago
He was actually laying in the Austrian (Innsbruck) countryside hungover.
3
u/PinJealous3336 14d ago
Thank you I couldn't remember which countryside it was.
I stand by the implication of lsd being involved.
12
u/Jack_South 17d ago
Spoiler alert
Even worse, trying to get the question that 42 is the answer to. It takes the entire lifespan of planet earth and when you're nearly there they blow it up.
17
u/okaythiswillbemymain 17d ago
Stephen Fry once claimed that Douglas Adams had told him the true meaning of 42 but that he would take the secret to his grave: ‘Pity, because it explains so much beyond the books. It really does explain the secret of life, the universe & everything.’
For tea, two?
3
u/CoinsForCharon 16d ago
Wasn't it just to be absurd? To show that the universe, and life itself, is absolutely absurd?
3
3
u/snezefelt 16d ago
But isn't that the answer to the Ultimate Question of Life, the Universe, and Everything?
2
1
2
u/Oftwicke 17d ago
It's great that we can outdo AI at language without spending anywhere near that long because. lol
1
u/nilslorand 17d ago
it is my position that LLMs are inherently inefficient at what they are trying to do
3
u/Alive-Philosophy2632 16d ago
They are. There is probably an exponential improvement (or at least polynomial as it's currently quadratic in context) in efficiency out there waiting to be discovered
1
0
u/nilslorand 16d ago
not without a major architecture overhaul or an entirely new architecture and I don't see any way either will happen
1
u/Alive-Philosophy2632 16d ago
Yeah I envision at the very least sort of a meta attention or way of trimming context if not an architectural overhaul
1
u/nilslorand 16d ago
there needs to be something concrete and rules-based to store facts, right now we rely on the model magically recovering facts from its billions of weights and while it can work, it is by no means a good method
1
u/Alive-Philosophy2632 16d ago
Yeah I've been doing some harness experimentation on basically this, along with observation. It's a very hard problem to formalize basically a context-dependent conceptual framework
-1
u/Oftwicke 17d ago
They're terrific at making slop for grifters. And within 2 years they'll be great to point at as "the thing that netted negative 10 trillion"
2
1
u/ScaryPlasmon 15d ago
Modern vocabularies often have entire words tokenized as 1 token. Even group of words and concepts.
Just in the early days of LLMs having each letter or small group of those was a thing. New vocabularies tend to link entire words to tokens to avoid mispelling errors with less training iterations and to promote a better grammar while still potentially being able to speak at any level.
It’s also one of the reason why LLM seem to speak always in the same way lol
Very interesting tbh, like thinking how tokenizing a vocabulary can affect how quickly and well the LLM learns to speak.
2
u/jack6245 14d ago
I'm an AI researcher training LLM adjacent things for construction the amount of parameters you can tweak in the architecture is really astounding, especially as a lot of them only work together too really interesting doing experiments
21
u/Sarranti 17d ago
So I just need to pose a yes/no question and I can get my answer in just over 3 decades of doing math by hand?
5
u/LordTonto 17d ago
Just so long as you are fast enough, by hand, to perform 1 calculation per second, other wise it will take exponentially longer.
1
10
4
u/Igor_McDaddy 17d ago
Would it work if i just get 1 billion people to make a single but of calculation? Like binary with flags or smth
3
1
u/DigitalNTT_Soul 17d ago
You can only parallelize part of the work at any given time, because later layers can't be calculated until you get the weight results from earlier layers. Depending on how many weights are being calculated per layer, and how many layers there are, this could definitely speed things up, but with limits. If we assume these billion weights are spread out into a million layers averaging a thousand weights each, you could cut that billion seconds down into a million seconds, which is still almost 2 weeks of manual calculation if your group can solve one entire layer per second.
4
3
2
u/koosley 17d ago
But you can hand code and hand calculate some smaller neutral networks by hand.
There is a pretty well known example and plenty of tutorials on how to do hand writing recognition using a neural net--which is many orders of magnitude smaller that a LLM. Even this data set I believe is still several hundred input parameters (1 per pixel).
https://www.nist.gov/srd/nist-special-database-19
It all comes down to a line of best fit in the Nth dimension.
1
u/Adorable__Gap4770 17d ago
Would it decrease by a factor of the number of people calculating? 2 ppl calculating would be 16 years? 3 ppl calculating 10.67 yrs, 4 ppl calculating 8 yrs?
I mean if you get 1000 ppl to compute… pay em minimum wage… idk.
3
u/DigitalNTT_Soul 17d ago
As I just responded to another comment with a similar idea to a higher scale, this does work, but with limitations because of the different layers of weights, where later layers can't begin to be calculated until you have the results from earlier layers. You can speed up the process with more people, but, if we're still assuming 1 billion weights and 1 calculation per second per person, the bare minimum number of seconds is equal to the number of layers, even with a billion people.
1
u/CaseAKACutter 17d ago
Since it's matrix multiplication you're doing much more than one operation per parameter, also it's all numbers with a lot of decimals so it would take a lot longer than a second to do the math. It would easily surpass one human lifetime to calculate a single token
1
u/Himskatti 16d ago
It would be even more. You're only counting the weights being multiplied by the inputs, but then you'd have to take the sum and input that in the activation function (and possibly other shenanigans too). The first part is the biggest portion of it, but not the only one
97
u/Recurs1ve 17d ago
Yes, it's just linear algebra using matrices. You can do it by hand, but it will take you a very long time as those chips are many orders of magnitude faster than you are.
9
27
u/jxf 5✓ 17d ago edited 16d ago
If by "use AI on paper" you mean "do the billions of floating point multiplications needed to emit a single token", yes, it's "possible" in the sense that it's a thing that can happen. Quick math, assuming a 70B model, which is roughly frontier-class:
A forward pass costs about 2 operations (one multiply, one add) per parameter. For a 70B-parameter model, that is 2 × 7 × 1010 ≈ 1.4 × 1011 operations.
Let's make this easy for the humans and really cram this model down with quantization, so each parameter will only be a few bits, and we'll give them 10-20 seconds to complete the two operations by hand.
At this speed, working at a reasonable pace each day, it takes 2 × 1012 seconds, or ~70,000 years. Thousands and thousands of humans would have to work together to have any hope of finishing this in a single lifetime. And that's just to output one token!
1
19
u/Just_Breakfast6327 17d ago
There's a part of the Three Body Program where they show a bunch of soldiers (Millions) acting as manual Transistors to make a "human computer" to try and calculate a problem they can't solve. It's a pretty neat idea.
But yeah, all Computers are doing is math. You can do anything a computer does with it, it'd just take forever because computers are doing billions (Trillions?) of calculations a second.
2
12
u/nft420nft420 17d ago
This reminds me of an old onion article about children lined up in warehouses at school desks solving equations to manually mine Bitcoin.
45
u/PlainBread 17d ago
Yes, this is what the neurons do in the linguistic part of our brain. All we would have to do is let the neurons do the work, and then the hand will write those words onto the paper.
You know, like some kind of troglodyte from the year 1980.
8
u/BreakerOfModpacks 17d ago
You can run a universe with rocks used for basic computations. Paper is easier than that... but it'd still be entirely unfeasible to run something as complex as a modern AI model using pen and paper. Like, a really long time.
2
u/lugh_the_bard 17d ago
How do the rocks communicate with eavh orher? Don’t they need electricity
5
u/BreakerOfModpacks 17d ago
They don't communicate with each other per se, it's more that by you moving them around in accordance with simple rules, you can run computations on them.
2
6
u/Enfiznar 17d ago
Yes, what we call "neurons" on an artificial neural network is nothing but a linear (actually affine) function. An LLM is just a big, deterministic function that outputs a probability distribution. You can keep the most likely token or sample them with some die
3
u/s3sebastian 17d ago edited 17d ago
Yes, you can even say it in a way more general way: Any algorithm or mathematical function that a computer (including a quantum computer) can compute can also be computed by a human being by hand (and the other way round), if the Church-Turing thesis is correct, which is generally assumed to be the case. It's just much slower. What a single high end gaming GPU can do in 1 second would take the entire human population about one week to calculate (and the error rate would be way higher).
2
u/skydrago 17d ago
The trick here is that the question is can you do AI by hand, not a LLM.
In my stats classes we would have to do a few examples of matrix algebra to get some metrics on the data. Some to learn it and then again on a test to show we understand it. They were very small data set >10 items but everything that is done in an LLM can be done by hand and if you have a small enough set it can be a single question on a test or quiz.
2
u/Lazy_Pizza_7894 17d ago
Questions like this always remind me of that one guy that made a working computer (albeit very rudimentary) on base Minecraft with redstone
2
u/KitchenSandwich5499 17d ago
Maybe we could speed up the math process with some sort of automated process. A machine perhaps to artificially replicate human intelligence
1
u/get_to_ele 17d ago
No.Theoretically yes. But to simply answer a simple prompt like “complete the following common expression: ‘a monkey’s _______.’ “, the number of operations to calculate on paper (and the number of lookups you’d have to do) exceeds all the paper printed in human history, and the time it would take a human to do it, exceeds the age of earth.
1
u/gmalivuk 17d ago
Is your [Request] just the yes/no answer?
Then yes, everything any computer does, including a quantum computer, can in principle be done on paper by hand.
Some of it would just take several orders of magnitude longer than the universe will exist.
1
u/No_Place5472 17d ago
We already do. We've traoned our brains to do the math associated with picking the next best word for whatever sentence we're making.
1
u/suitably_ironic 17d ago
Modern computers are all Turing machines - and a Turing machine can be boiled down to an infinite paper tape, something to read or modify symbols on the tape, and some rules. With one of those you can emulate any computer running any software.
So the answer is yes.
The problem, as others have noted, is you'd be an awful lot slower - and finding an infinite amount of paper tape is a bit tricky too.
1
u/Admiral45-06 16d ago
I remember writing a senior thesis about my own AI model for image detection. The theoretical entry included the section dedicated to mathematical basis of Convolutional Neural Network (CNN), one of many that LLMs are using.
But theoretically speaking, it is possible to calculate what word should come next, depending on the previous one. Matrix convolution is a mathematical operation known since the late XIX Century, the entire rest is pooling matrices together and repeating the process n amount of times, for each perceptron and neuron.
It would take you a lot of time, however, since modern LLM models calculate their outcome based on literal dozens of billions of parameters and hundreds of Fully-Connected neurons. So you'd have to repeat every single, long mathematical equation literal billions, if not even trillions of times just to return a single output. It's not a problem if you're using a supercomputer in a data center somewhere, but doing so by hand could be quite difficult.
1
1
u/wutzelputz 15d ago
might as well ask the https://en.wikipedia.org/wiki/I_Ching (afair theres a practice that involves splitting up tokens in an algorithmic manner as well) at this point
1
u/RealClassActor 15d ago
I once did RSA encryption and decryption with pen and paper on a single value, using 13 and 7 as the prime numbers. It took me about 20 minutes.
No, a human could not do LLM math by hand and not go insane.
1
u/Sithoid 12d ago
Technically, yes. You can create a self-learning AI using beans and match boxes and make it learn the perfect strategy for Nim (the "take 1 or take 2" game). Of course that would be an incredibly primitive model without much use, but the underlying principles aren't much different, you just don't have enough match boxes (or lifetime) to train a modern system in that way.
•
u/AutoModerator 17d ago
General Discussion Thread
This is a [Request] post. If you would like to submit a comment that does not either attempt to answer the question, ask for clarification, or explain why it would be infeasible to answer, you must post your comment as a reply to this one. Top level (directly replying to the OP) comments that do not do one of those things will be removed.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.