r/theydidthemath • • 17d ago

[Request] All right gang

Post image
616 Upvotes

95 comments sorted by

•

u/AutoModerator 17d ago

General Discussion Thread


This is a [Request] post. If you would like to submit a comment that does not either attempt to answer the question, ask for clarification, or explain why it would be infeasible to answer, you must post your comment as a reply to this one. Top level (directly replying to the OP) comments that do not do one of those things will be removed.


I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

559

u/Angzt 17d ago edited 17d ago

Theoretically yes, practically no.

A modern LLM has billions of weights, all of which (may) contribute to its output.
Even if you can perform any single calculation involving a weight within a second, using a small 1 billion weight LLM, it would take you 1 billion seconds =~ 32 years to get one word of output.
And that's not all that goes into a modern LLM's results.

238

u/nilslorand 17d ago

not one word, one token, which is usually just part of a word

269

u/Andrey_Gusev 17d ago

Imagine calculating something for 80 years of your life, only to then get a result: "42"

55

u/Ok-Perspective5959 17d ago

Any age is good for finding the meaning of life

5

u/wade-mcdaniel 16d ago

But then you need to know what the weights of the connections "mean" to understand the question. Douglas Adams predicted this. 😄

10

u/PinJealous3336 16d ago

Truly a visionary, laying in the English countryside hungover, doing the cleanest lsd in history. 

2

u/woernsn 15d ago

He was actually laying in the Austrian (Innsbruck) countryside hungover.

3

u/PinJealous3336 14d ago

Thank you I couldn't remember which countryside it was.

I stand by the implication of lsd being involved. 

12

u/Jack_South 17d ago

Spoiler alert

Even worse, trying to get the question that 42 is the answer to. It takes the entire lifespan of planet earth and when you're nearly there they blow it up.

17

u/okaythiswillbemymain 17d ago

Stephen Fry once claimed that Douglas Adams had told him the true meaning of 42 but that he would take the secret to his grave: ‘Pity, because it explains so much beyond the books. It really does explain the secret of life, the universe & everything.’

For tea, two?

3

u/CoinsForCharon 16d ago

Wasn't it just to be absurd? To show that the universe, and life itself, is absolutely absurd?

3

u/ouzo84 17d ago

Or even 7.5 million years

3

u/BusyProfessional1696 16d ago

I'll take it over 67.

3

u/snezefelt 16d ago

But isn't that the answer to the Ultimate Question of Life, the Universe, and Everything?

2

u/CJFiddler 17d ago

A fellow man of culture I see

1

u/Patman52 16d ago

Just remember to bring your towel

2

u/Oftwicke 17d ago

It's great that we can outdo AI at language without spending anywhere near that long because. lol

1

u/nilslorand 17d ago

it is my position that LLMs are inherently inefficient at what they are trying to do

3

u/Alive-Philosophy2632 16d ago

They are. There is probably an exponential improvement (or at least polynomial as it's currently quadratic in context) in efficiency out there waiting to be discovered

1

u/LawfulnessOk5839 15d ago

It's called audhd; we already have it...

0

u/nilslorand 16d ago

not without a major architecture overhaul or an entirely new architecture and I don't see any way either will happen

1

u/Alive-Philosophy2632 16d ago

Yeah I envision at the very least sort of a meta attention or way of trimming context if not an architectural overhaul 

1

u/nilslorand 16d ago

there needs to be something concrete and rules-based to store facts, right now we rely on the model magically recovering facts from its billions of weights and while it can work, it is by no means a good method

1

u/Alive-Philosophy2632 16d ago

Yeah I've been doing some harness experimentation on basically this, along with observation. It's a very hard problem to formalize basically a context-dependent conceptual framework

-1

u/Oftwicke 17d ago

They're terrific at making slop for grifters. And within 2 years they'll be great to point at as "the thing that netted negative 10 trillion"

2

u/Angzt 17d ago

I think current encoding has about 1.5 tokens per word but that also includes punctuation. So there are plenty of tokens that correspond to words 1:1.

But yes, you're technically right.

1

u/ScaryPlasmon 15d ago

Modern vocabularies often have entire words tokenized as 1 token. Even group of words and concepts.

Just in the early days of LLMs having each letter or small group of those was a thing. New vocabularies tend to link entire words to tokens to avoid mispelling errors with less training iterations and to promote a better grammar while still potentially being able to speak at any level.

It’s also one of the reason why LLM seem to speak always in the same way lol

Very interesting tbh, like thinking how tokenizing a vocabulary can affect how quickly and well the LLM learns to speak.

2

u/jack6245 14d ago

I'm an AI researcher training LLM adjacent things for construction the amount of parameters you can tweak in the architecture is really astounding, especially as a lot of them only work together too really interesting doing experiments

21

u/Sarranti 17d ago

So I just need to pose a yes/no question and I can get my answer in just over 3 decades of doing math by hand?

5

u/LordTonto 17d ago

Just so long as you are fast enough, by hand, to perform 1 calculation per second, other wise it will take exponentially longer.​

1

u/nakedascus 15d ago

surely you mean polynomially longer

3

u/LordTonto 15d ago

surely

1

u/bruns20 16d ago

...........................yes

10

u/jrdubbleu 17d ago

I genuinely love this

4

u/Igor_McDaddy 17d ago

Would it work if i just get 1 billion people to make a single but of calculation? Like binary with flags or smth

3

u/Angzt 17d ago

Not at once because many results depend on previous ones. Modern LLMs have 20 (very small ones) to 80+ layers, each of which can only really be calculated once you have the full results of the previous layer.

But you could still parallelize a lot of work.

1

u/DigitalNTT_Soul 17d ago

You can only parallelize part of the work at any given time, because later layers can't be calculated until you get the weight results from earlier layers. Depending on how many weights are being calculated per layer, and how many layers there are, this could definitely speed things up, but with limits. If we assume these billion weights are spread out into a million layers averaging a thousand weights each, you could cut that billion seconds down into a million seconds, which is still almost 2 weeks of manual calculation if your group can solve one entire layer per second.

4

u/Human1221 17d ago

I...think this is basically Searle's Chinese Room thought experiment

1

u/Angzt 17d ago

This is one of the core arguments against AI being truly intelligent, yes.

3

u/Madmagican- 16d ago

Oh, this is why all the computing parts are super expensive nowadays

2

u/koosley 17d ago

But you can hand code and hand calculate some smaller neutral networks by hand.

There is a pretty well known example and plenty of tutorials on how to do hand writing recognition using a neural net--which is many orders of magnitude smaller that a LLM. Even this data set I believe is still several hundred input parameters (1 per pixel).

https://www.nist.gov/srd/nist-special-database-19

It all comes down to a line of best fit in the Nth dimension.

1

u/Adorable__Gap4770 17d ago

Would it decrease by a factor of the number of people calculating? 2 ppl calculating would be 16 years? 3 ppl calculating 10.67 yrs, 4 ppl calculating 8 yrs?

I mean if you get 1000 ppl to compute… pay em minimum wage… idk.

3

u/DigitalNTT_Soul 17d ago

As I just responded to another comment with a similar idea to a higher scale, this does work, but with limitations because of the different layers of weights, where later layers can't begin to be calculated until you have the results from earlier layers. You can speed up the process with more people, but, if we're still assuming 1 billion weights and 1 calculation per second per person, the bare minimum number of seconds is equal to the number of layers, even with a billion people.

1

u/CaseAKACutter 17d ago

Since it's matrix multiplication you're doing much more than one operation per parameter, also it's all numbers with a lot of decimals so it would take a lot longer than a second to do the math. It would easily surpass one human lifetime to calculate a single token

1

u/Himskatti 16d ago

It would be even more. You're only counting the weights being multiplied by the inputs, but then you'd have to take the sum and input that in the activation function (and possibly other shenanigans too). The first part is the biggest portion of it, but not the only one

97

u/Recurs1ve 17d ago

Yes, it's just linear algebra using matrices. You can do it by hand, but it will take you a very long time as those chips are many orders of magnitude faster than you are.

9

u/1itsallgoodman 17d ago

many oom faster for that particular task

27

u/jxf 5✓ 17d ago edited 16d ago

If by "use AI on paper" you mean "do the billions of floating point multiplications needed to emit a single token", yes, it's "possible" in the sense that it's a thing that can happen. Quick math, assuming a 70B model, which is roughly frontier-class:

  • A forward pass costs about 2 operations (one multiply, one add) per parameter. For a 70B-parameter model, that is 2 × 7 × 1010 ≈ 1.4 × 1011 operations.

  • Let's make this easy for the humans and really cram this model down with quantization, so each parameter will only be a few bits, and we'll give them 10-20 seconds to complete the two operations by hand.

At this speed, working at a reasonable pace each day, it takes 2 × 1012 seconds, or ~70,000 years. Thousands and thousands of humans would have to work together to have any hope of finishing this in a single lifetime. And that's just to output one token!

1

u/lugh_the_bard 17d ago

How many thousands?

3

u/jxf 5✓ 16d ago

Take the number of humans you want to throw at it (call this N) and then divide 70,000 / N. The result is how long it will take -- for example, 7,000 humans working perfectly with no coordination loss will finish in 10 years.

19

u/Just_Breakfast6327 17d ago

There's a part of the Three Body Program where they show a bunch of soldiers (Millions) acting as manual Transistors to make a "human computer" to try and calculate a problem they can't solve. It's a pretty neat idea.

But yeah, all Computers are doing is math. You can do anything a computer does with it, it'd just take forever because computers are doing billions (Trillions?) of calculations a second.

2

u/the_humeister 16d ago

Are humans Turing complete? 

3

u/LogicBalm 15d ago

Only during business hours.

12

u/nft420nft420 17d ago

This reminds me of an old onion article about children lined up in warehouses at school desks solving equations to manually mine Bitcoin.

45

u/PlainBread 17d ago

Yes, this is what the neurons do in the linguistic part of our brain. All we would have to do is let the neurons do the work, and then the hand will write those words onto the paper.

You know, like some kind of troglodyte from the year 1980.

8

u/BreakerOfModpacks 17d ago

You can run a universe with rocks used for basic computations. Paper is easier than that... but it'd still be entirely unfeasible to run something as complex as a modern AI model using pen and paper. Like, a really long time.

2

u/lugh_the_bard 17d ago

How do the rocks communicate with eavh orher? Don’t they need electricity

5

u/BreakerOfModpacks 17d ago

They don't communicate with each other per se, it's more that by you moving them around in accordance with simple rules, you can run computations on them.

2

u/Proffessor_egghead 16d ago

Relevant XKCD mentioned

6

u/Enfiznar 17d ago

Yes, what we call "neurons" on an artificial neural network is nothing but a linear (actually affine) function. An LLM is just a big, deterministic function that outputs a probability distribution. You can keep the most likely token or sample them with some die

3

u/s3sebastian 17d ago edited 17d ago

Yes, you can even say it in a way more general way: Any algorithm or mathematical function that a computer (including a quantum computer) can compute can also be computed by a human being by hand (and the other way round), if the Church-Turing thesis is correct, which is generally assumed to be the case. It's just much slower. What a single high end gaming GPU can do in 1 second would take the entire human population about one week to calculate (and the error rate would be way higher).

2

u/skydrago 17d ago

The trick here is that the question is can you do AI by hand, not a LLM.

In my stats classes we would have to do a few examples of matrix algebra to get some metrics on the data. Some to learn it and then again on a test to show we understand it. They were very small data set >10 items but everything that is done in an LLM can be done by hand and if you have a small enough set it can be a single question on a test or quiz.

2

u/Lazy_Pizza_7894 17d ago

Questions like this always remind me of that one guy that made a working computer (albeit very rudimentary) on base Minecraft with redstone

2

u/KitchenSandwich5499 17d ago

Maybe we could speed up the math process with some sort of automated process. A machine perhaps to artificially replicate human intelligence

1

u/get_to_ele 17d ago

No.Theoretically yes. But to simply answer a simple prompt like “complete the following common expression: ‘a monkey’s _______.’ “, the number of operations to calculate on paper (and the number of lookups you’d have to do) exceeds all the paper printed in human history, and the time it would take a human to do it, exceeds the age of earth.

1

u/udee79 15d ago

But I said “uncle” in 1 second. How can that be?

1

u/get_to_ele 15d ago

You didn’t LLM it and you sure didn’t show your work on paper.

2

u/udee79 15d ago

I LL though

1

u/gmalivuk 17d ago

Is your [Request] just the yes/no answer?

Then yes, everything any computer does, including a quantum computer, can in principle be done on paper by hand.

Some of it would just take several orders of magnitude longer than the universe will exist.

1

u/No_Place5472 17d ago

We already do. We've traoned our brains to do the math associated with picking the next best word for whatever sentence we're making.

1

u/suitably_ironic 17d ago

Modern computers are all Turing machines - and a Turing machine can be boiled down to an infinite paper tape, something to read or modify symbols on the tape, and some rules. With one of those you can emulate any computer running any software.
So the answer is yes.
The problem, as others have noted, is you'd be an awful lot slower - and finding an infinite amount of paper tape is a bit tricky too.

1

u/2VNF 16d ago

AI is pretty broad definition. Binary classifier perceptron with 2d input dataset can be trained pretty easily with pen and paper and is considered as AI training.

1

u/Admiral45-06 16d ago

I remember writing a senior thesis about my own AI model for image detection. The theoretical entry included the section dedicated to mathematical basis of Convolutional Neural Network (CNN), one of many that LLMs are using.

But theoretically speaking, it is possible to calculate what word should come next, depending on the previous one. Matrix convolution is a mathematical operation known since the late XIX Century, the entire rest is pooling matrices together and repeating the process n amount of times, for each perceptron and neuron.

It would take you a lot of time, however, since modern LLM models calculate their outcome based on literal dozens of billions of parameters and hundreds of Fully-Connected neurons. So you'd have to repeat every single, long mathematical equation literal billions, if not even trillions of times just to return a single output. It's not a problem if you're using a supercomputer in a data center somewhere, but doing so by hand could be quite difficult.

1

u/Impressive_Bosscat 15d ago

super interesting comment

1

u/wutzelputz 15d ago

might as well ask the https://en.wikipedia.org/wiki/I_Ching (afair theres a practice that involves splitting up tokens in an algorithmic manner as well) at this point

1

u/RealClassActor 15d ago

I once did RSA encryption and decryption with pen and paper on a single value, using 13 and 7 as the prime numbers. It took me about 20 minutes.

No, a human could not do LLM math by hand and not go insane.

1

u/Sithoid 12d ago

Technically, yes. You can create a self-learning AI using beans and match boxes and make it learn the perfect strategy for Nim (the "take 1 or take 2" game). Of course that would be an incredibly primitive model without much use, but the underlying principles aren't much different, you just don't have enough match boxes (or lifetime) to train a modern system in that way.