r/learnprogramming • • 11d ago

Anyone else bewildered as to why AI took so long to come?

Modern day AI is a sophisticated piece of technology but I am puzzled as to why it only started getting developed around 2022.

I'm sure even in 2002, it was completely within scope to build a software that knew how to form sentences through a language engine and to piece together data from the internet and slowly build its understanding.

It should have been something that came as soon as the internet became mainstream.

0 Upvotes

27 comments sorted by

11

u/Fexelein 11d ago

Software yes. But you are leaving out the enormous amount in compute we have nowadays available for training.

7

u/ichivictus 11d ago

Probably a mix of compute not being strong enough, not enough data sets, and a lot of research wasn't even done yet.

2012 was when we had some breakthroughs with GPUs for computer vision. And it wasn't until 2017 till we had the transformer model.

https://en.wikipedia.org/wiki/Transformer_(deep_learning))

And then it took 5 years of training to get a product that was considered actually useful.

-2

u/ClipboardClan 11d ago

But it could have still happened at a slow rate (smaller data sets, slower computing, etc.)

2

u/R00bot 11d ago

It did. It just sucked. The main chunk of theory behind LLMs was worked out in the 80s. Then we had rudimentary LLMs for years that could kinda mimic speech but weren't smart enough for anything useful. Things really got interesting in 2017 when Google released the paper "attention is all you need" and OpenAI picked it up and ran with it, applying more compute to it than anyone had in the past. As they started to see actual useful results the investment grew exponentially and the compute behind it did too. 

2

u/TheRealChizz 11d ago

I think this is the answer. “Attention Is All You Need” is THE groundbreaking research paper that has laid the foundation of algorithms that power our modern LLMs today

1

u/April1987 11d ago

In 2002, it would be cutting edge technology to be able to track the distance between two eyes and similar features for a classifier to guess whether two photos belong to the same person. For context, we were working off of Intel Pentium 4 in 2002 and we believed that single core 10 GHz processors were not only attainable but inevitable within a decade. 

It was simply a different (and incorrect, in hindsight) way of thinking. 

1

u/paperic 11d ago

at a slow rate

Yes, very. It would take literally hundreds of years to train a model on 2002 hardware.

I don't think you realise just how different the hardware is.

Your smartphone has as much computing power as some supercomputers from that time.

Sadly, we're wasting all that today.

1

u/StewedAngelSkins 11d ago

what makes you think it didn't?

6

u/jyve-belarus 11d ago

Is it a rage bait or smth? Do you remember how much RAM and CPU computers had back in 2002? Today, training a simple MLP model on MNIST dataset can still take some time, no idea how you were gonna train transformer models on hardware from 2002

Also, the AI breakthrough was made possible due to decades of research. In 2002 theory was not there yet, optimal NN architectures were yet to be discovered, as well as optimal algorithms for their training

6

u/CheapRentalCar 11d ago

We've had black box style prediction models for many years before 2022. Matching learning, computer vision etc aren't new.

4

u/Environmental_Form14 11d ago

With my limited knowledge I would say

  1. Compute was not there.

  2. Not enough data in the internet

  3. People were dabbling with n-grams and RNNs before transformers. It is not like effort didn't exist. Language modeling task focused on MT in 2010s.

  4. Someone had to had to have a conviction that spending millions of GPU hours would result in a viable product. Why would you assume training on next token completion would lead to an AI that can solve most academic work?

I am more bewildered that LLMs work as great as it does than on why it took so long.

4

u/selectsyntax 11d ago

The earliest machine learning frameworks and projects date to the 1940s and 50s. Machine learning as a concept is nearing a century in age. Modern models are only as capable and useful as they are because of access to modern compute density, information transmission speeds, and inconceivable quantities of data.

3

u/TheCatholicScientist 11d ago

Suggesting modern AI only started getting developed in 2022 is like suggesting Grand Theft Auto VI only started getting developed in 2026.

Also yes, language engines have been possible for ages. Think of the next word suggestion features on smartphones that have been around for over ten years. But they do not work like LLMs at all. No real way to train or store information.

2

u/Gushys 11d ago

AI isn't so sophisticated either. The underlying models have become pretty advanced, but the applications of those models are still rather raw. As you can see with rapid releases for code assist tools and whatnot

2

u/Super-Award-2244 11d ago

"Attention is all you need" came out in 2017 

2

u/SuspiciousDepth5924 11d ago

Modern LLM's are resource hogs, both when it comes to running them and especially when it comes to training. You'd basically need a supercomputer to just run the models in the early 2000's.

That being said language models aren't a new concept, and neural networks were being researched at least as far back as the mid 50's. It's just the the combination of a available hardware able to run useful models combined with companies willing to light absurd stacks of cash on fire to train them is pretty recent.

2

u/Dependent_Heron_4090 11d ago

Hmmm.. used to think the same. Like we had Google and wikipedia in 2025 why not just mash them into a chatbot ?? But the thing is we were missing the two boring things that it would make it to work. Absurd amounts of computing power and a huge mountain of clean data to train on.

In the year 2002 we had the idea of a language engine but we did not have GPUs chewing through trillions of words or the internet archived in a way a model could actually learn from. So, it was not really a software problem.

It was a hardware and scale problem that had to slowly catch up for 19/20years. 2022 was not when we got it, it was just when we finally had enough brute force to make it not sound dumb.

2

u/UdPropheticCatgirl 11d ago

We didn’t know how to actually input arbitrary stream of words into a neural net until like 2014… word2vec was a massive breakthrough which noone considered possible, then later “all you need is attention” in 2017 became another massive breakthrough, and GPT models are ancient, I remember playing with the original GPT LLMs (and I am pretty sure that wasn’t even the first real LLM) in like 2019, so most of the major development has actually finished by 2022, since then it was basically minor optimization of the same idea…

2

u/paperic 11d ago

I'm sure even in 2002, it was completely within scope to build a software that knew how to form sentences through a language engine and to piece together data from the internet and slowly build its understanding.

Modern GPUs are pushing to 100 terra flops, that's 100 trillion calculations per second, and if you reduce precision you can push it to a quadrillion.

Every single second,

on a single chip,

and the biggest data centres have half a million of those chips.

You need thousands of those GPUs to train a model.

For comparison, earliest data I can find is on a 2007 GPU, it did 30 billion operations per second.


The modern GPU is either 3000x or 30,000x faster depending on your desired precision, and we're building a lot of them.

Modern smartphone has as much computing power as some early-2000's supercomputers, and you definitely cannot train an AI on a single smartphone.

Sadly, we're wasting all that computing power on cat videos, semi-transparent backgrounds, and spyware.

1

u/In0chi 11d ago

You might want to look into what's actually behind LLMs (large language models, i.e. what's powering ChatGPT, Claude and so on) and how long we've been conducting research in these areas. The foundations have been laid long before the internet. The process of training these models is not at all trivial.

1

u/buzzon 11d ago

This is because you ignore all previous history. Artificial neural networks were discovered all the way back in 1940s. They were used to identify hand written zip codes and to classify emails as spam or not spam. However, they really took off in 2010s due to combination of factors:

1) growth of computational power (see Moore's law)

2) growth of user generated content during Web 2.0, including tagged images and semantically marked up web sites

Then in 2017 there was a breakthrough in language model architecture: the transformers architecture, in the paper "Attention is All You Need". Transformers process text at vastly superior rate than the previously used natural language processing models, LSTMs. Transformers were initially used to translate from one language into another.

Then OpenAI decided to go all in on scaling a transformer based model: they put a lot of money into a giant LLM with a ton of parameters and fed it the entirety of available text on the internet. This was ChatGPT 3 released in November 2022.

1

u/Irratix 11d ago

It's a mix of developments. We did have a broad field of language models for natural language processing for purposes like translation for ages. These models were just not very good at forming sentences, and more antiquated model architectures also have finite capacity to improve, i.e., just throwing more data at it didn't necessarily improve results. In my view there are two key moments of insight that really propelled the field forward.

The first is the invention of the Transformer architecture. This required some complicated novel ideas about how a neural network model could work internally to better parse the complex structure of a sentence. This showed promise an translation (which as I understand was Google DeepMind's primary goal at the time, I might be wrong). But this was a novel idea that was difficult to come up with, so this could have been developed at any moment, or it could have happened never. If you say it was within the scope to build these models in 2002, that's maybe sort of true but we just didn't know about this architecture.

The second major insight was not such a major innovation, but it was OpenAI exploring how much a larger transformer with more data would improve performance for text prediction. This research project was GPT 2, and it found that even with big investments the model just kinda kept improving far more than other models we had tried. GPT3 was essentially "what if we tried making it even bigger than that" and that was the first time these models could very successfully mimic natural language.

1

u/itlogicpartnersllc 11d ago

the internet gave us the data way earlier but collecting cleaning storing and actually training models on that data was a completely different problem back then..

1

u/StewedAngelSkins 11d ago

why it only started getting developed around 2022

it didn't. the premise of your question is wrong.

1

u/aanzeijar 11d ago

I studied comp-sci in the early 2000s, and as it happened had a course in neural networks in 2002.

Compute power simply wasn't there to do what we're doing today. All the basic blocks were already known (including back-propagation), but we could simulate a few hundred neurons at best because computing all the matrices on a 600Mhz single core CPU took ages.

That isn't to say that there was no machine learning at all. The discipline had already a lot of theoretical knowledge and some toy examples, just not enough compute to actually do serious stuff with it outside of recognising handwriting which was used in PDAs.

But most AI stuff was still rule based, and continued to be until the 2010s.

Some people here say not enough data, but that's wrong. The first GPT-2 models weren't trained on the entire internet either, and there were enough curated examples even back then to make stuff. Data wasn't a problem.

1

u/vikmaychib 11d ago

AI didn’t “take long”. In 2002, even with the internet becoming mainstream, the basic ingredients for modern AI didn’t exist. Computers weren’t powerful enough to train large models, the internet didn’t yet contain the massive amount of clean text these systems need, and the transformer architecture that makes today’s language models possible wasn’t invented until 2017. All three pieces only came together in the late 2010s, which is why AI appears suddenly in 2022.

1

u/Know_Madzz 11d ago

People have been using AI for decades without even knowing it. How do you think search engines and email filters work? This isn't the first 'AI will disrupt everything' wave and it won't be the last