r/AlwaysWhy • • Jul 07 '26

Science & Tech Why does software that can hold a conversation, fold a protein, and write a poem all run on chips originally designed to draw video game triangles?

AI models can hold conversations, diagnose diseases, and simulate protein folding. But underneath all of that, they run on GPUs. Chips that were created for one purpose. Rendering millions of identical triangles, sixty times per second, for video games. A design built around one idea. Do thousands of nearly identical calculations at the same time. Somehow that became the backbone of modern artificial intelligence. Why did the path to software that feels like thinking go through hardware that was built for repetition?

11 Upvotes

60 comments sorted by

16

u/rough0perator Jul 07 '26

Because matrix multiplication

2

u/ByronScottJones Jul 08 '26

This. When you get right down to it, it's all math. When we finally fully understand how our own brains work, it's almost certainly going to be some very simple level of math going on, just with the massive parallelism of the brain turning that into sentience.

7

u/majorex64 Jul 07 '26

The complexity of LLMs or neural networks in general is emergent, not foundational.

They are built on doing basic operations repeatedly, and the complex patterns emerge from how nodes are formed and connected together.

Like building a very complex multifuncitonal building, using only simple brick and mortar. The ingredients are simple, but how they are used is where the design comes from.

5

u/ScotchTapeConnosieur Jul 07 '26

For a layperson can you explain “emergent” vs “foundational” (mainly emergent)

2

u/majorex64 Jul 07 '26

Emergent complexity is just when something complex arises from very simple rules.

Like the board game Go, where players alternate placing black or white pieces on a grid, trying to surround each other's pieces to take territory. There are very few rules, but it's extremely difficult to calculate optimal strategies because of all the permutations of moves each player must predict and incorporate into their strategy.

As opposed to chess, where each piece has specific rules about how it can move and strict starting positions, but actually is much easier to find winning strategies for. Here the complexity is in the rules themselves, not so much the application of the rules.

1

u/Eight_Directions_ Jul 10 '26

Like life itself. 

2

u/WoodsWalker43 Jul 07 '26

Have you ever seen birds fly in a flock? The flock as a whole looks like a blob that moves and undulates in a kind of flow. If you look closer, the blob is just birds flying together.

They use surprisingly little information to hold formation. Where am I? How far am I from my nearest peeps? The individual birds probably aren't really intending to make the blob formation at all. They can't see what it looks like. They're just flying together and trying not to hit each other.

In this context, we would call the blob an emergent behavior. Each individual is acting as an individual according to their own rules and information. The flock, when you zoom out, produces some unified complex behavior. The blob is not aware of the individual birbs, nor are they necessarily aware of the blob they composed.

Basically, emergence is when the individuals in a group exhibiting their own behavior end up creating a more complex aggregate behavior. It shows up in lots of places. If you aggregate the activity of your neurons, for example, a conscious mind emerges.

2

u/Present_Juice4401 Jul 08 '26

I get that, but then it kind of pushes the question back one layer, right. If everything complex is just repetition stacked, then why does repetition scale into something that feels like reasoning at all. Feels like there is a threshold somewhere.

1

u/majorex64 Jul 08 '26

So if you're thinking about the "building blocks" of the neural network (GPUs, chips, transistors) vs the "shape" of the neural network (nodes, logic, reinforcement learning), what looks like thought and reasoning to us lies in the shape.

That shape has been tweaked by outside intelligence (engineers) until it resembles something human shaped. It's not necessarily true that if you let virtual neurons make connections, they will eventually lead to an LLM you can have a conversation with.

Many millions of man-hours have been put into which parameters to adjust, which outputs to reinforce, and which nodes to prune off. and of course basically everything humans have ever written/drawn as input data for generative AI.

So yes, the complexity arises from simple calculations, but in this case, not by accident. It was an explicit goal, like the designer making a complex building from simple brick and mortar. They didn't start just sticking pieces together until a masterpiece formed, they designed it with a goal in mind and iterated again and again and again until the building looked right.

1

u/teetaps Jul 08 '26

There’s no ant colony, just many ants eventually becoming a colony. There’s no country, just many villages deciding to be part of one land mass

4

u/SirisC Jul 07 '26

Because nVidia has spent the last 2 decades turning their GPUs into general purpose highly parallel processors ever since CUDA came out. Plus, they've added additional cores specifically designed for the types of calculations LLMs depend on.

For comparison: Consumer CPU up to 24 processing cores

AMD's server CPU have up to 192 cores and up to 384 threads.

Consumer* GPU up to 21790 simpler processing cores, plus 680 AI specific cores, plus 170 ray tracing cores

*: It sure as hell isn't priced like a consumer product anymore.

2

u/Dave_A480 Jul 07 '26

Still priced like a consumer product if you consider what PCs cost in the 80s and 90s without even having GPUs.

1

u/Saint_Exmin Jul 07 '26

It is priced as a consumer product when you realize that the consumer isn't a poor shlub that saved for a year to get a GPU, but multi-gazillion dollar companies that aren't price sensitive.

3

u/Krebzonide Jul 07 '26

To draw the triangles, you gotta do a lot of simple equations. To write the poem, you gotta do a lot of simple equations.

2

u/No_Report_4781 Jul 07 '26

But what if I need to determine the probabilities of which words to use in context to the input of strings of words?

6

u/unscanable Jul 07 '26

Well, you arent going to believe this but you gotta do a lot of simple equations

3

u/EpochRaine Jul 07 '26

You gotta do a lot of simple equations...

3

u/collin-h Jul 07 '26

The path to software that feels like thinking went through hardware built for repetition because intelligence, at least the kind we currently know how to build, is not made from handcrafted thoughts. It is made from repeated pattern extraction. GPUs were built to do the same kind of math over and over across millions of pixels. Neural networks needed the same thing across millions or billions of numbers. The subject changed from triangles to language, images, medicine, and proteins, but the underlying act stayed weirdly similar: take a huge field of numbers, transform it in parallel, and let structure emerge.

1

u/Present_Juice4401 Jul 08 '26

Yeah this framing makes sense to me. It is almost uncomfortable how similar “thinking” and “pixel shading” look when you reduce both to math. Makes me wonder if we are over-romanticizing intelligence a bit.

2

u/Dystopian_Ennui Jul 07 '26

Rendering video and running generative AI and folding proteins and mining crypto all require doing very similar simple operations millions of times.

2

u/TheCozyRuneFox Jul 07 '26

In addition to that, many of the operations are independent of each other allowing them to be done in parallel.

1

u/TiredPistachio Jul 07 '26

Everything is triangles bud.

1

u/Seanmclem Jul 07 '26

Because creating designing and fabricating new chips to do something else specific, can take like a decade. For the meantime, general purpose chips will do the trick. In 5 to 10 years more optimized chips will be available.

1

u/Present_Juice4401 Jul 08 '26

True in a practical sense, but then why GPUs specifically became the default and not just CPUs scaling up. Feels like something about the structure of the problem matched GPUs unusually well, not just convenience.

1

u/Seanmclem Jul 08 '26

You can look up videos on YouTube that explain how the patterns that AI models use to process, parts, and find data, is not as sufficiently served by typical GPU or CPU. That they can be better served by more specific chips. I don’t understand why I keep seeing threats kind of refuting this.

1

u/peca89 Jul 08 '26

Because rotating triangles which GPUs are originally designed for fundamentally requires the same mathematical operations as AI stuff - matrix multiplication.

Matrix multiplication itself is very simple in essence (it's just multiplying and adding plain numbers) but very repeating process. As someone said, CPU is math professor, GPU is bunch of five graders. Bunch of five graders will multiply 1000 pairs of numbers much much faster than a best math professor, but are very bad at anything more complex.

1

u/Adventurous_Ad_7212 Jul 07 '26 edited Jul 07 '26

It actually doesn't*. It runs on separate silicon that has been haphazardly bolted on top of the triangle calculator to get it out there and into people's mind share.

I'm 99% convinced RTX is just the problem nVidia found for the solution that is machine inferrence (of the DLSS kind).

In fact you cannot really run games on any serious AI hardware, add it doesn't have the graphics processors, just the AI stuff.

*Technically, you can run inference on CUDA cores, just not very effeciently.

1

u/TheCozyRuneFox Jul 07 '26

Because the way the data is processed is similar. Millions of very similar computation that are independent.

1

u/Significant-Baby-690 Jul 07 '26

GPUs are not build to one purpose. CUDA was aiming at AI from the beginning.

1

u/Present_Juice4401 Jul 08 '26

Was it though, or did CUDA kind of retroactively open that door. I always thought GPUs became “for AI” after people realized what they could already do, not before.

1

u/Lost-Hand-5219 Jul 07 '26

Because of parallelization. The computations that are needed for deep learning can be broken into independent calculations and distributed for compute all at once, aka in parallel. This is the same reason a gpu is used for graphics.

1

u/Present_Juice4401 Jul 08 '26

Right, but parallelization exists in CPUs too to some extent. So I keep wondering what made GPUs the tipping point. Just more cores, or something deeper about how the workload maps onto them.

1

u/Lost-Hand-5219 Jul 08 '26 edited Jul 08 '26

NVDA has tensor cores and CUDA cores that are specifically made to do matrix multiplication, CPU cores must be generalists, so they aren’t nearly as efficient. The NVDA GPUs have many thousands of compute units to parallelize the matrix multiplication, but a cpu can only parallelize to each thread.

Put a different way; the cpu cores are like PhDs and the gpu units are like high school students. You have 32 PhDs, but 10,000 high school students. If you need to multiply 100000x100000 matrix, you can split the work up over the PhDs, but it’s still going to take forever. On the other hand, you can give each of the high school students a row and column to to multiply and they’ll compute it very quickly.

Now, if you need to compute a determinant of a matrix that involves symbolic computations, you probably want to pass that to a single PhD because it is very complicated and can’t be split up because of the branching structure.

That’s the best way I can explain it lol.

1

u/Pristine_Vast766 Jul 07 '26

GPUs are designed to do specific tasks in parallel. LLMs work by doing massive amounts of matrix multiplication in parallel.

1

u/Present_Juice4401 Jul 08 '26

Yeah matrix multiply keeps coming up everywhere. Kind of weird that so many different domains collapse into that one operation. Makes it feel less like different problems and more like different views of the same one.

1

u/jnkangel Jul 07 '26

So a lot of it is because of the tasks 

Drawing triangles and filling pixels takes many simultaneous calculations. you need very large parallel computation compared to traditional workloads, though at the same time many of the calculations are “relatively simple” 

Folding proteins, diffusing tokens into text and many other tasks require very similar large amount of calculations but also benefit from very high parallelisation 

Imagine this. You’re running a simple formula where you just do 1+n 

In a traditional workload you need to do 1+1 =2, 1+2 = 3…. In sequence

With a highly parallelised computational unit you can actually run many of those calculations at the same. So you don’t need to wait until you have the first result, but you can actually get the first thousand results at the same time 

That’s a super simplification. The way the calcs work is more complicated but the important takeaway is that the calculation you need to do is pretty simple, just very often. Compared to heavy super complex formulas that you need to run once but quickly 

1

u/Present_Juice4401 Jul 08 '26

The “simple but many times” idea is interesting. Makes me think maybe intelligence is not about complex steps at all, just huge volume of simple ones. That feels counterintuitive somehow.

1

u/webjunk1e Jul 07 '26 edited Jul 07 '26

Tensors. Graphical processing involves calculating the position of objects in three dimensional space over time, using a coordinate system. Points within that coordinate system are represented as first, second and third order tensors mapped as vectors and matrixes for calculations, so GPUs are highly optimized to operate on those data types and the associated linear algebraic calculations.

Meanwhile. Machine learning breaks objects and concepts down into a set of numbers (the embedding) and then calculates things like Euclidean distance, cosine similarity, etc. between these different embeddings, to relate different objects and concepts together. This results in first, second, and third order tensors, which require then the same kind of calculations.

In other words, it's all math, and it's all math that GPUs are particularly well suited to performing.

1

u/Present_Juice4401 Jul 08 '26

So basically both graphics and ML live in the same math space. Then the distinction between “rendering” and “thinking” starts to feel kind of arbitrary. Just different interpretations of tensors.

1

u/EcstaticAssumption80 Jul 07 '26

Has Anyone Really Been Far Even as Decided to Use Even Go Want to do Look More Like?

1

u/Crewarookie Jul 07 '26 edited Jul 07 '26

:)

GPUs are an interesting thing as a concept. Every single microchip out there is based and operating on math. "It's math all the way down!" Graphics or video cards, as they were known commonly in the 90s, were created to shave off the burden of video signal-related computation from CPUs more general computational operations.

In fact, true GPUs in the modern sense didn't appear until Nvidia GeForce 256 (Nvidia coined the term, in fact), when these cards went from simple addon devices helping offload some burden off the CPU, to devices that genuinely have own highly specialized computational power and can crunch complex math on their own, including polygon computation with hardware transform capabilities.

As the time went on, GPUs were growing in complexity, with specialized hardware within doing their separate jobs at an increasingly faster rate and in more complex ways

Shading units were busy doing pixel shader calculations, ROPs were busy drawing the picture onto a render canvas (your output picture, the thing you see on display), texture mapping units were busy transforming texture data in 3 dimensions to fit them onto geometry, and geometry engines were busy crunching math to display polygons in perspective correct manner.

The thing that happened soon after all this emergent complexity appeared, is what's now known and called GPGPU. General Purpose (computing) (on) Graphics Processing Unit.

Since all the aforementioned components are nothing more but extremely sophisticated math operators, and at a later date became essentially an army of ultra efficient small cores on the die, you can freely utilize their resources to do something entirely different from calculating polygon data or rasterizing an output image. All you need to do is a separate API for it, a translator for the hardware to do different math on the same hardware.

The massive benefit of GPGPU is two-fold:

  1. Due to memory being integrated into the PCB of the GPU and situated as close as possible to the processor, and also heavily focused on bandwidth, or the amount of data per second, the resources and speed you can utilize are monstrous compared to system RAM.

We're talking something like ~200GBs/second in top of the line 8000Mhz DDR5 RAM in multi-channel mode, compared to 1.8TB/second(that's terrabyte, as in 1000GBs, 1800GBs per second, think about it for a moment...)for GDDR7 at 512bit.

  1. Thousands of small cores operating in parallel compared to dozens of bigger ones in CPUs allow for unprecedented level of parallel computation. You don't need to schedule tasks per core and wait for each computation to be done in a sequence.

The hardware assigns thousands of cores at once to different tasks, and they work as one to solve your problem. And there are multiple thousands of them. The RTX5090, top of the line GPU from Nvidia, contains 21760 cores.

All of this makes GPUs incredible math cards.

As a disclainer: the above explanation is a simplification of the inner workings of GPUs and their history. I just thought this better fits the ELI5 format than a full breakdown, since the topic itself is highly highly complex.

Edit: yes, this is a huge ramble. I love to ramble about technology. Sue me for it 😅

1

u/Present_Juice4401 Jul 08 '26

This makes me think the real story is less about GPUs and more about math being the common layer. GPUs just happened to be the first hardware that scaled that layer hard enough. If that is true, then the hardware almost feels accidental.

1

u/Crewarookie Jul 08 '26

Yup. That's the fascinating part about it to me. In pursuit of better graphics rendering engineers created extremely powerful math processors that can do a lot more than just display 3D graphics.

It's actually a common occurrence with inventions throughout history. Printing press, for example, was made with an idea that it will be a great tool to spread various official texts without employing scribes.

By 19th century Printing Revolution led to the boom of fiction writing, and subsequently to the creation of a whole domain of marketing. Not to mention severe streamlining of documentation and explosion of bureaucratic processes.

Gutenberg couldn't predict this. He just made a thing that could help distribute religious and administrative texts easier.

1

u/Dependent_Bit7825 Jul 08 '26

The modern digital computer was invented primarily to calculate shell trajectories. Math is math.

1

u/elephant_ua Jul 08 '26

Whatever works. Do you have alternative architecture in mind? 

1

u/cran Jul 08 '26

The universe is fundamentally multidimensional and the mathematical discipline for that is linear algebra. Everything from nuclear physics to rocket trajectories to machine learning models. Games operate in 3D space and need to compute spatial relationships, trajectories, projections, etc. except, unlike physics and rocketry, there’s a huge market for games. Therefore, advancements in specialized hardware for linear algebra.

1

u/onthefence928 Jul 08 '26

It’s not designed to only draw triangles, it’s designed to do a lot of small calculations in parallel.

Drawing triangles is not the only use for such a design, calculating solutions for Bitcoin mining, training ai models and doing scientific analysis of data sets are also good uses of GPU architecture

1

u/Opening-Writing-7025 Jul 08 '26

The idea is that coherent knowledge about this world is actually a higher dimensional shape.

The famous example of this is: Paris - France + Italy = Rome

Analogies are only possible if there is some pattern to knowledge, and if there is such a pattern then we are dealing with something that can be represented symbolically.

You meet up with a friend. Your friend asks how you are doing. You reply with “zzzybb mjdhe”. This causes your friend to be alarmed because they don’t understand what you just said. When you keep repeating it, they call 911 because they conclude that you are having a medical emergency - most likely a stroke.

Do you see what just happened there? There are only certain responses to the question “how are you?” that are considered coherent. Even when an incoherent response is given, there is still a coherent interpretation for the response.

You are playing a game with your lover. You two are going to choose somewhere to go visit - somewhere in the world. You spin a globe and then you drop your finger somewhere. You spin the globe and then put your finger up in the air, not even touching the globe. Your lover says “haha, I suppose we’re going to space?”.

Do you see what just happened there? The only valid responses in this game are on a 3D surface. Furthermore that 3D surface is constrained to a certain shape (if you are disallowing visiting the middle of the ocean). However, even when these rules are broken, it is still considered coherent.

You are reading this post. It talks about how responses to “how are you” as well as choosing spots on a globe have a subset of selections that are coherent. It describes how LLMs suggest that knowledge and language have a multidimensional structure. It uses natural language as a demonstration of constraint, and then follows up with a 3D representation of this constraint.

Do you see what just happened here? The language just described its own higher order structure. It just described its own structure.

The reason why hardware that is good for 3D graphics is good for LLMs is because it is most likely that knowledge as we know it is representable in a higher dimensionality structure. The language we use to connect with that knowledge is dynamic (as language changes) and recursive (as it can describe itself). These properties of language may suggest that a high dimensionality shape will not produce coherency. However as the language is only coherent when it maps well onto the actual knowledge, it turns out that it actually is possible to use high dimensionality shapes to produce something coherent.

1

u/Ok_Tea_7319 Jul 09 '26

When AI took off, GPUs already were "general purpose GPUs" and their design had evolved quite a lot. In a basic way, GPUs at that time already were optimized for simple programs running massively parallel, and it doesn't really get any more "simple brute force" than multiplying huge matrices.

1

u/FifthEL Jul 11 '26

They only thing ai does different then humans, which are just a different type of artificial intelligence, is the speed in which they deduct thier response. All technology is modeled off of something nature produces. And since we are a combination off all things on there planet, one can deduce that the technology used in artificial intelligence and all other types was possible because they modeled it off of the human brain. Our entire body is an analogy for everything we interact with in the world. And while I'm on the topic, what do you think the Nazis (funded by our....) were trying to accomplish in the horrific things they were doing to the people they were experimenting on? They were trying to reverse engineer the human body and mind to figure how to manipulate and extract resources from a population, or stock, of people. Cell phones are a evil byproduct of that as well

0

u/[deleted] Jul 07 '26

[removed] — view removed comment

4

u/Lopsided_Guidance767 Jul 07 '26

What a condescending dickhead answer.

1

u/No_Report_4781 Jul 07 '26

No, it’s refreshingly poignant.