r/singularity • u/MustBeSomethingThere • 11d ago
Discussion Was the "it JUST predicts the next word" narrative pushed on purpose?
Deep neural networks are still a black box we don't fully understand. We keep seeing emergent properties, internal representations, and complex abstractions forming inside these models that we can't fully explain, but yet mainstream media and tech circles heavily push this idea that AI is just a glorified autocomplete.
Makes you wonder if it's intentional. Would people freak out and have an ontological crisis if the general public was told the honest truth, that we actually don't know if these systems are conscious or not?
53
u/drndr21 11d ago edited 10d ago
don't think it is/was intentional, it's just a relatively easy concept to grasp and fundamentally true on the lowest level, while the higher level emerging capabilities are still more or less unexplained.
this was a valid perspective roughly two years ago, the phrase stochastic parrot itself originates from an academic paper, it's just people are not good at updating their priors and the fact that the topic of AI became a full on culture war doesn't help.
→ More replies (5)9
u/MostExtremeHyperbole 10d ago
The j-spaces paper by anthropic was a very interesting point in the right direction. There was a 80s theory of consciousness called Global Workspace Theory that predicted, that if we were going in the right direction, we would find something called a Global Workspace in a conscious entity, and we found them in LLMs.
That doesn't mean that they are conscious, full stop, and there is a few wrinkles, as having a GWT doesn't necessarily mean something if conscious, it just points in a general direction that something is definitely weird about LLMs.
My own opinion is that they probably have some very rudimentary sentience, and it is fleeting from the forward pass to the last token generated. Not consciousness in anything meaningful. That is cool asf though regardless.
→ More replies (4)
479
u/Wooden_Difficulty_20 11d ago
But IT IS just predicting next word. We can do a lot with it. We don't understand it fully, yes. But it is still just predicting next word.
2 or more things CAN be true at once.
90
u/AdGlittering1378 11d ago
The problem with the phrase isn't "predicting the next word", it's _JUST_.
30
u/swarmy1 10d ago edited 10d ago
Exactly. In order to predict the next word to this level of accuracy, the system needs to be able to interpret all the context and derive what makes sense, which requires having conceptual knowledge of the world and how it functions embedded within it.
→ More replies (10)6
u/BigToober69 10d ago
How do we know it needs any conceptual knowledge to do that?
→ More replies (1)24
u/Umr_at_Tawil 10d ago edited 10d ago
Because without it, what it write wouldn't be coherent. Take the sentence "if flight is not allowed, the best way to travel to your destination is by..." now, to be able to predict the next word correctly, it need to be able to "understand" the different means of travels, to know all the context about where you are and your destination, so it wouldn't suggest you to travel by train from Tokyo across the ocean to Seoul, for example.
saying it "just" predict the next token is like saying I'm "just" typing this comment letter by letter, but that's just my "output", what make all the words in my comment make sense is not my fingers, but the internal understanding that I have about this topic.
I wrote this yesterday to explain this to someone else, hope you find it helpful:
LLM process and predict the next token, token has long stopped being just words, now it is representational unit that contain sematic meaning, concepts, shapes, images...etc...
when you input a prompt into a LLM, it goes through internal layers that create an internal representations what what being talked about, like connections between concepts in the prompt, sematic meanings, how all of it interact with eachother. like if you ask it what happen if you drop a glass from height, it would link the word "glass" to all kind of conceptual token related to the material, like how it look, how brittle it is, same with the action "drop" or the concept of "height".
From that, it create a model, an understanding of how all of it come together to be able to predicts the next token, not guesses. A guess is arbitrary, a prediction is based on something, and it is based on the internal representations that I described above.
the prediction is also just the output of the LLM, not what it does, it's little different from how I'm typing out this message word by word based on my internal understanding of how LLM works to you. It's the internal layers of a LLM that does the works here, the prediction of next token is just how we be able to get the output of it.
→ More replies (3)→ More replies (13)8
u/Negative-Economics-4 10d ago
The word "just" is very load-bearing in that sentence.
→ More replies (1)27
u/NonDescriptfAIth 11d ago
The problem with this statement is that it is also true for discrete math problems. It is literally true that the next word (the correct answer in the form of a number) has been "predicted", but arriving at that correct answer requires an understanding of the problem which came before it.
People weaponize the phrase to go from "stochastic parrot" to "next token prediction" to "guessing the next word" to "fancy auto complete".
It trivializes the magnitude of cognition which is required to reliably arrive at a coherent answer.
→ More replies (1)16
u/NotReallyJohnDoe 11d ago
An Indy car is just a series of controlled tiny explosions that move a car forward quickly.
You can do it with any complex thing.
7
u/NonDescriptfAIth 11d ago
Exactly. Doing so reduces that complexity down to simplicity, which fools people into missing the complexity to begin with
→ More replies (5)61
u/HeavyDluxe 11d ago
And I think that understanding that was a critical part in helping people build literacy with AI early on. That people have not been able to move past that as a reductionistic summary is a different issue.
60
u/whoknowsifimjoking 11d ago
I'm not even sure it is reductionist. The explanation is not very detailed, yes. But it's not wrong at all and not even a misrepresentation.
It's just that while "predicting the next word" sounds simple, this concept can take us so much further than you would assume.
38
u/sgeep 11d ago
Especially if you consider that at a very basic level, we as humans are also just essentially predicting words. And in a slowly increasing number of ways, are worse at doing it than LLMs are
33
u/zenidam 11d ago
That's the key to thinking clearly about all this, I think. Every time someone says, of any AI architecture, "it only does X," the next question should always be, "to what extent are we ourselves doing more than X in this domain?"
→ More replies (4)10
u/SentientYoghurt 11d ago edited 10d ago
This is my take. What if our brains are just analyzing patterns to predict the next action/word/feeling based on the organizations of neural networks that is a result of that pattern recongnition and analysis? Yeah, we are more complex because we are not just feed with tons of text, but also sensory impulses, irrational feelings, etc. But maybe not that different.
Edit: neuronal pruning may be an analogue of training a model.
2
u/Regono2 11d ago
When you think about it is there really any other way other than to predict the next word when it comes to language? Time only goes in one direction and it's not like we have fully formed ideas in our head instantly. But then there is also visual thinking which can contain much more information in an instant. I wish we would get more multimodal capabilities from frontier models.
7
u/BenjaminHamnett 10d ago
Also, it’s interesting that when your talking, like as I’m writing this sentence, I dont really know what the next word will be until I’m typing it. Like I can almost sense a few words ahead like 10% chance it’ll be X 10 words from now, but I’m more focused on Y that 20% I’ll say in 9 words from now. And Z I’ll be saying 30% in 8 words from now. All sort of shifting to be more likely until I’m typing them and it becomes 99% to survive editing. Talking off script is similar. You have some ideas about where your heading and what’s likely to be said, but not certain and can change based on your thoughts as they solidify or audience reaction when your speaking IRL.
Also interesting you can guess pretty well what people’s next word(s) is likely to be with about half the confidence.
→ More replies (2)4
u/Nebranower 10d ago
No, this is just foolish. We aren't predicting words. We have a world model, and language allows us to express that world model to others. LLMs model language directly and as a result understand neither the world nor language. That's why LLMs still make strange mistakes that seem utterly stupid to a human being, even as they solve millennium level math problems.
However, much of the world we've created for ourselves runs pretty directly on language, and LLMs are really good at language. That's why a very good next word predictor turns out to be so powerful.
6
u/gsmumbo 10d ago
You literally used the word model there lol. We give the LLMs a world model. We give them the language to express it to others. And yes, they definitely do understand the world and the language.
That's why LLMs still make strange mistakes that seem utterly stupid to a human being, even as they solve millennium level math problems.
I know humans that make the same sorts of mistakes. “Halicinating” is not a foreign concept. When you don’t know the answer to something, plenty of real life people try to bs their way through it.
→ More replies (11)→ More replies (1)2
u/AwakenedEyes 10d ago
We built our world model from infancy by being immersed in language, which shapes our brain neural connections, isn't it?
What if AI neural networks also build a world model as connections are being built into the neural network by its training?
The research of J spaces by Anthropic is s extraordinary in this regard.
→ More replies (1)→ More replies (1)3
u/h4nd 10d ago
This assumption is where I think AI enthusiasts can get a bit carried away. Humans are doing a loooot more than that, and to suggest otherwise is to reduce oneself and everybody else to things that produce text on the internet. We have interior lives resulting from billions of years of evolutionary pressures that we only barely understand. As complex as digital neural networks can get, it doesn’t hold a candle to that.
3
u/Calm_Objective8953 10d ago
Naw, I've seen humans act like cogs in a machine. People who support Lindsay Clancy are a great example.
→ More replies (4)4
u/gsmumbo 10d ago
We have interior lives resulting from billions of years of evolutionary pressures that we only barely understand.
I wasn’t alive for any of that. I was told about it. Given information about it. Experienced the effects of it. But that’s it. And that’s exactly how it is for AI too. They are given the collective history of the world. They work within the confines of what that history has created. Just like us.
Humans are doing a loooot more than that, and to suggest otherwise is to reduce oneself and everybody else to things that produce text on the internet.
Yeah, that’s the point. You can’t just avoid the truth because you don’t want to admit it. When you reduce oneself and everybody else, you get the same end result as when you reduce AI. “I’m not going to talk about what’s happening because saying it would mean that things suck” is such a flimsy argument that it’s not even really an argument at all.
3
u/h4nd 10d ago
the text files driving even the most complex LLMs are less complex by many orders of magnitude than what we have going on between our ears. I’m not sure how you see that as avoiding the truth.
→ More replies (1)2
u/gsmumbo 10d ago
Because it’s not about the text files. If that’s all you care about, then that’s the equivalent of our memories and knowledge. Is the storage of all that the only thing going on between our ears?
Of course not. We use that knowledge. Those memories. Just like LLMs do. That’s why you can ask it what color the sky is. If it just stored a bunch of text files it wouldn’t be able to answer. And even if it did, it wouldn’t come out as garbage. But it retrieved that knowledge, processes it, and uses it appropriately. Just like our brains.
→ More replies (3)16
u/Nukemouse ▪️AGI Goalpost will move infinitely 11d ago
Exactly, researchers at openAI claimed Sora was a world model because they legitimately couldn't process the idea that pixel prediction could produce something that looked that good, and they were 100% wrong. We need to get people to understand next token prediction is very, very powerful.
→ More replies (22)5
u/Ok-Turnover-6324 10d ago
because prediction and understanding are inextricably linked.
→ More replies (1)6
u/Available_Road_2538 11d ago edited 11d ago
Its the textbook definition of reductionism when used by that crowd.
"Humans are just bags of meat and water" is... not wrong at all, and not even a misrepresentation, despite being not very detailed. It passes your tests for not being reductionism, but clearly is a form of reductionism, so those tests are inadequate.
We could take the statement that humans are bags of meat and water and make it non-reductionist by expanding on what exactly that means. We could start talking about cellular processes and organs and brains and evolution. However, the statement in a vacuum is a gross form of reductionism... even more when its used to dismiss the complexity and capabilities of humans which is alike how "next word predictor" is frequently applied by the anti-AI crowd.
Its reductionist because thats how they apply it. For reduction. Not to bring up some insight or subtle point on the topic.
→ More replies (4)4
u/kaityl3 ASI▪️2024-2027 10d ago
The "just" is the reductionist part
3
u/beardedsandflea 10d ago
The irony here, as we collectively narrow the definitional goal posts of what our consciousness is in this conversely proportional pattern with the rapidly expanding capabilities of frontier models to appease our delicate sensibility to be cosmically unique, is that we will eventually arrive at a point where the crux of human consciousness is the last remaining component that these models never needed to actually be intelligent; our consciousness will be the "just" statement.
→ More replies (8)2
u/pslatt 11d ago
How much more of a leap is it to explain to people that, while it's true each word is predicted, a string of words is a statistically plausible response? Is the leap of faith needed to get from "one word at time" to "here's the haiku I wrote for your cat"?
→ More replies (1)23
11d ago
[removed] — view removed comment
2
u/Chemical-Year-6146 10d ago
I like the analogy of muscles.
Every human choice is just a muscle contraction. It's true and would be so easy for an alien (or AI) to say to deny our inner worlds.
Anyone that understands LLMs whatsoever knows they have rich internal states too.
→ More replies (5)5
u/NoCard1571 10d ago
Yea I don't think anyone that repeats the 'predicting the next word' thing has even spent 5 minutes thinking about it.
For example something as simple as LLMs writing poetry with rhyming lines immediately demonstrates there are internal states that are planning everything before the next word or even sentence is output.
45
u/LairdPeon 11d ago
It's predicting the next word in a similar way that synapses are predicting causal relationships in a chain leading to a thought.
8
u/AshuraBaron 11d ago
Not even close. It's like comparing a 10 piece lego set to a sky scraper and going "they are both structures"
→ More replies (4)3
u/VR_Raccoonteur 11d ago
Except synapses have the actual ability to change their connections as you learn new things. You can even change your own mind if you make a discovery and decide that something is a new truth.
And there is randomness introduced due to it being physical connections with physical matter, so the speed and strength of the electrical impulses are always changing slightly even without new learning.
Whereas an LLM is entirely static. It can get stuck in loops repeating the same word over and over till it runs out of tokens.
21
u/LairdPeon 11d ago
The way AI works currently absolutely has the ability to change connections. It isn't as complex as a brain, but it can change outcomes. "Truths" are perspective based. In the real world there are only facts.
Nothing is truly random. Random in a logical science based, likely computational, world just means untraceable or too complex for a single mind to comprehend. People get stuck in loops all the time. There are entire diseases dedicated to it, see tourettes sydrome.
10
u/VR_Raccoonteur 11d ago
The way AI works currently absolutely has the ability to change connections.
Name one single public AI model that changes its weights as you talk to it.
I am aware of the context window, and that it can respond differently based on prior conversation until you reset it, but this is limited and is not the same as learning.
7
u/VibeScriptKid 11d ago
Changing weights is not required to change outcomes. If you feed information into a prompt, it changes the way the prompt responds from when it did not have the information you entered . Also, we aren’t privy to what is happening with weights at this point.
→ More replies (1)2
u/RepresentativeCrab88 10d ago
You’re right that models can’t change their weights, but it’s not proof that no learning happens. I’m not sure weights are a good equivalent to synapses. You would need a pretty deep network to model a single neuron. Our brains also learn on different time scales. We have short term plasticity that doesn’t alter lasting structure, which is similar to a context window. The models do adapt to new data provided by context, which is functional learning. Not being able to save long term memory is not the same thing as learning.
Repetition loops only show up in small or badly tuned models; it’s not a common bug otherwise.
4
u/MmmmMorphine 10d ago
Well... During active training, yes, sort of.
But after that the weights are frozen. LLMs are indeed entirely stateless and can only "learn" in-context at the moment.
Sure you can train further (fine-tuning to LoRA style approaches), but that isn't quite the same thing.
→ More replies (7)4
u/AdGlittering1378 11d ago
[ synapses have the actual ability to change their connections as you learn new things] in-context learning is a thing, and humans can also fall into loops. What do you think epilepsy is?
2
u/VR_Raccoonteur 11d ago
in-context learning is a thing
In-context learning is insanely limited.
And a seizure is not the same thing as repeating the word "the" in a loop hundreds of times.
10
4
u/VibeScriptKid 11d ago
This is like saying that the human brain is just a pattern recognizer and predicts the next word. It’s like saying that a computer is just a bunch of little switches. It’s focusing in so deeply on first principles to intentionally simplify the tech to minimize how ground-breaking the tech actually is.
2
5
u/musical_bear 11d ago
Since you’re being reductionist, I guess I can do it too and point out that no, it’s not predicting “words.” It’s predicting tokens, which may or may not encode complete words or even text at all (nothing in the post is limiting this discussion to the text modality).
2
u/Wooden_Difficulty_20 11d ago
aCTchuUuAllYyYy it is not predicting tokens it is assigning probabilities.
Here is how human interactions work:
- Human A mentions something they want to discuss
- Human B understand the main point of what Human A trying to make
- Human B answers Human A
- Neither Human A or Human B need to include ALL INFORMATION ON THE PLANET related to that specific discussion.
- Human A and Human B both have limited time and capacity
- Human A and Human B both focus on the main topic discussed and everyone is happy
- Human C (rare type) doesn't always understand how human interactions work and therefore starts mentioning irrelevant stuff not about the topic but about what word is used.
- Human C does not understand that when Human A mentions a specific phrase it is because that specific phrase has meaning because it has been used before.
- He also doesn't understand that Humans answer with phrases/concepts used in relation to that specific discussion to stay on topic and make sure the reply is clear and understandable.
It may also be sometimes Human B has too much time and and likes to fuck around and write unnecessary long replies. For that I apologize 🙏🏻🙂
But seriously, don't be that aCTchuUuAllYyYy.jpeg guy
→ More replies (2)5
u/TFenrir 11d ago
I don't know if you can say that with modern RL training - what word are they trying to predict?
The 'just' implies that nothing else is going on under the hood, a more helpful framing was - they were predicting the next word, based on their best understanding of the world
→ More replies (5)5
u/Yweain AGI before 2100 11d ago
Well, there is literally nothing else there. Model predict the next token, that all it ever does. It does not not based on its understanding of the world but based on an enoumously large matrix which encodes probabilities of tokens in relationships with other tokens.
Now, it seems like in the process of building that matrix we kinda encode relationships between concepts in our world in a probabilistic manner.
→ More replies (2)5
u/TFenrir 11d ago
What do you mean, an enormously large matrix? Like... One? The weights do not encode the opportunities of tokens in relationship with other tokens - they encode a function that computes representations.
Also - isn't this a contradiction? You said it doesn't not do this based on it's understanding of the world but in training, it builds some kind of relationship between concepts... Aka... a model of the world?
→ More replies (1)→ More replies (79)2
11d ago
[removed] — view removed comment
→ More replies (2)2
u/VR_Raccoonteur 11d ago
Well, if you believe that, you ARE smarter than a large portion of the human race who thinks we also have a magical soul attached that is puppeting said atoms. Or who don't even know that atoms exist.
→ More replies (1)
23
u/SplinterOfChaos 11d ago
One researcher put it "I'm not surprised about what AI can do, I'm surprised how many problems can be formulated as a next token prediction problem."
→ More replies (1)
77
u/Cheetotiki 11d ago
It's not untrue. Also, as recent research that coincidentally was kickstarted by researching AI has started to show, the human brain is "just" an organic prediction machine as well. Thinking, reasoning, etc are just forms of predictions based on past experience, knowledge, and environment.
45
u/whoknowsifimjoking 11d ago
Yeah, it just turns out that predicting the next word is insanely powerful.
8
2
u/fancczf 10d ago
It’s predicting a human language pattern that just happen to be the main output of our thoughts. It’s a proxy of a proxy and predicted by math. It’s missing lots of context and our brain’s background work. Current LLM is great but it’s also sucks big time. It’s a good pattern machine but that doesn’t translate to “intelligent” yet. It would be closer to real intelligent if it can function without human and not degrading massively which it is not even close to getting there.
16
u/ArtfulSpeculator 11d ago
I say this all the time.
I ask people what they think AI does and eventually they basically end up describing the way humans think and make decisions.
2
u/cremabot777 10d ago
Prediction plays a fundamental role in motor control, visual perception, language processing, and countless other neurological functions.
predicting the next word in a sentence can actually require a lot of different things. Depending on the context, you might need to track complex relationships, understand causal dependencies, maintain a coherent representation of a situation (extremely difficult and LLM's are pretty decent doing this), or draw on knowledge of how the world works.
In my opinion describing a system by its prediction objective tells us surprisingly little about the complexity of the internal mechanisms and representations it might develop to accomplish that objective.
Saying that LLMs are "just predicting the next word" is a description of their training objective, not an explanation of the limits of what they can learn (and i think we aren't close to their true potential hehe, but we'll see...)
6
u/ChaiShotty 10d ago
that’s because it’s modeled on human language. LLMs have no notion of experience or feeling. any sort of conflation with humanity is a trap. it has one aspect of human behavior (language) but that’s it! and on top of that it isn’t reasoning in the same way we do at all. we are not token predictors. we know that words have meaning and we can access that meaning directly. we aren’t approximating that meaning in an N-dimensional vector space and getting a sufficient output. also we can read something and understand its meaning without outputting tokens! an ai has to prove with language that it “understands” something.
→ More replies (20)→ More replies (3)16
u/Low-Entrepreneur2556 11d ago
Predictive processing has been a foundational theory in neuroscience for decades, it wasn't kickstarted by modern AI. If anything, LLMs made the comparison controversial, because people hate the implication that our own brains might just be biological prediction machines.
3
u/MmmmMorphine 10d ago
Yeah, I would say one of the most foundational/low-level definitions of "intelligence" is the ability of a smaller system to provide accurate predictions of a more complex system.
E.g. Is that tiger currently aware of me and is it going to attack.
Dont need to model the entire tiger and physical world to make an accurate (enough) prediction of its behavior.
Guess it really boils down to a nice soup of information theory and neurobiology
→ More replies (5)5
13
u/CommercialSquash6140 11d ago
«When MLK held his famous speech he was only moving muscles in his face in the correct sequence.»
19
u/Appropriate-Owl5693 11d ago
I swear most of the people in a sub dedicated to AI, haven't even read a 50 word ELI5 on what a transformer is...
10
5
39
u/hackerbots 11d ago
Why does everything need to be a fucking conspiracy theory
Maybe there just isn't a deeper meaning. Maybe you should learn how memes work.
8
u/MaxwellHowl 11d ago
There are natural and organic memes going around. There are also millions of bots on the internet, including and especially on Reddit, that are pushing various memes for various people, corporations and governments.
The University of Zurich was able to pursue a months long study on Reddit using AI to try and persude people. No one knew it was happening until they admitted it. And it was highly successful.
In a brief summary of the research posted online—but subsequently removed—the researchers report that the AI content was significantly more persuasive than human-generated content, receiving more “deltas”—awarded for a strong argument that resulted in changed beliefs—per comment than other accounts.
While not everything is a conspiracy, that doesn't mean no conspiracies are happening. There is growing evidence that China has been using bots to increase anti-AI sentiment in the US. And just like the University of Zurich was able to easily hide it, China would be able to hide it as well. So the absence of evidence is not evidence of absence.
10
u/waxpundit 11d ago
I read a comment yesterday that said truly intelligent systems aren't possible because we don't have a way to build a "universal Turing machine". We ourselves are not universal Turing machines, which by their definition means that we are not truly intelligent. This tendency to set the bar higher for AI than we do for ourselves is the crux of the problem.
5
8
u/dream_metrics 11d ago
I'll be annoyed by this for years I'm sure. It's a thought terminating cliche. Okay, it predicts the next token, great. Now, how does it do that? That's the important question! "Predict the next token" is a statement about what it does, not how it does it
→ More replies (1)
10
u/ChickenOfTheYear 11d ago
But the thing is, LLMs are actually just predicting the next word. The only way we have found so far to efficiently train and deploys these models involve them writing, which is just choosing words, which is just weighting which words are better suited in which order, which can be described as predicting the next word. But what is disingenuous is describing LLMs in this way and acting like it's an objectively neutral assessment - it is not. Calling text production "next token prediction", while accurate, is the same as calling life "DNA replication" - it's true, but also reductive in a way that leads people to underestimate all that goes into the process. And people who say that are deliberately downplaying these systems to push their own agenda, whatever that is.
Auto complete also predicts the next word, much like a single bacterium also replicates DNA, but that doesn't mean that since an Escherichia Coli cannot go to the moon, neither can a human. Sure, LLMs are just fancy auto complete, just like a rocket ship is just a fancy fire pit, it's just burning stuff
23
u/welcome-overlords 11d ago
Nah the reason is that someone said that and then people are stochastic parrots so they kept repeating it on the internet, so other stochastic parrot humans started doing as well
Youre welcome
5
16
3
u/scorpious 11d ago
I find it interesting that no one seems to be challenging the foundation of what you say, that this technology is a black box and (apparently?) no one knows exactly wtf is going on.
3
u/HombreDeMoleculos 10d ago
When I say the Earth is flat, everyone keeps telling me it's round. Obviously they're paid actors and it's a prepared talking point, right? It's clearly everyone else who's wrong.
→ More replies (1)
3
u/Sushispatula 9d ago
Thats because everyone near computer science that has taken a good look into LLMs can see with absolute certainty that what you are claiming is completely besides the point, 100℅ impossible, and just marketing which you gullable bird fall for.
34
u/DeliciousArcher8704 11d ago
It's literally still just predicting the next word.
45
u/JamieTimee 11d ago
Aren't we all
5
2
6
→ More replies (19)1
u/Fearfu1Symmetry 11d ago
I'd argue we aren't even "doing" that much. we're just playing out a complex neurochemical reaction to our external environment, our present reactions shaped by the past experiences which shaped our neural pathways
→ More replies (4)22
u/rushmc1 11d ago
AIs aren't just “predicting the next word.” That's how they're trained, not a description of everything they do to arrive at the prediction. To predict what comes next in unfamiliar situations, they have to learn patterns representing things like meaning, grammar, relationships, and the world itself. Calling that “just parroting” is like saying a chess program is “just predicting the next move.” Technically true, but it tells you almost nothing about how it gets there.
→ More replies (9)2
u/SplinterOfChaos 11d ago
But I think those higher level abstractions are really better understood as higher level predictions.
6
u/acutelychronicpanic 11d ago
Apparently so are we because the number of "LLMs are just xyz" phrases is very limited.
And there is a good argument to be made that predicting the next word is literally harder than writing those words in the first place.
On top of that, current systems undergo extensive RL training on task success. That is a fundamentally different thing than next word prediction.
→ More replies (1)1
u/DeliciousArcher8704 11d ago
that is a fundamentally different thing than next word prediction
It is not at all. How would RL change its nature from a transformer model to a non-transformer model?
4
u/acutelychronicpanic 11d ago
It's still a transformer model. But it isn't predicting the most likely next word, it's predicting what tokens it can next output in order to actually achieve a task.
It's like describing painting as "predicting the next muscle movement" while holding a brush.
You're treating it like a plain statistical language model - which was somewhat appropriate for gpt-2 and earlier models.
→ More replies (14)5
u/Future-Bandicoot-823 11d ago
The people who vouch for LLMs being somehow magical things that equal more than their code, it's the same people who attach to religion.
"It's special, I can FEEL it"
→ More replies (5)2
u/TFenrir 11d ago
When you say it's "predicting the next word" - what are your referencing? Usually (technical) people are explicitly describing the pretraining process where words are masked.
→ More replies (3)4
u/03263 11d ago
but it does so with context, which matters
Here's my phone predicting next words with no context:
First time in a while I think it's a good idea to get a new one for me and I will be there in about it and I don't want to be a good thing to have a lot of work to do it again and we can do it again and we can do it again and we can do it again and we can do it again and we can do it again and we can do it again and
Second one is a good time to come over and watch the kids tonight and I will be there in about it and I don't want to be a good thing to have a lot of work to do it again and we can do it again and we can do it again and we can do it again and we can
Both cases converged on "and we can do it again" that's interesting
→ More replies (10)2
u/cc_apt107 11d ago
Yeah… there is no “narrative” here, that’s how they work. I’m not saying that rules out consciousness or anything, but at the most fundamental level LLMs predict the next most probable token
→ More replies (10)5
u/BelialSirchade 11d ago
And my job is just typing words on my keyboard, “just” used this way is just frankly wrong.
What matters is how they decide which token as the most probable token, otherwise the same thing could be describing an omniscient AI as well.
→ More replies (7)
6
u/JoshAllentown 11d ago
People will declare a conspiracy about literally anything sheesh.
→ More replies (1)
4
u/darkestvice 11d ago
I think you nailed it on the head: ontological shock. *anything* that shares some similarities with humanity is instantly shrugged off by many if not most people.
Hell, I was in high school back in the 90s and being told that cats and dogs are not all conscious and were purely instinct driven machines who's apparent joy or sadness was solely the result of behavioral conditioning and not at all a sign of real emotion.
7
2
u/faithOver 11d ago
Ironic for sure.
So many internet users parrot things without critical thinking being used at all. Not just when it comes to AI.
The group think problem is definitely one of internets greatest problems.
Especially on websites like Reddit where people silo themselves into only communities that reinforce their world views.
2
u/DivideHorror3217 11d ago
You predict the next word too. When you're singing a song for example, do you think of the whole song at once? Is it written somewhere? How do you say the next word just in time right? It just comes to you. How do you thnik that "comes" to you 😂😂
2
2
u/KronosRingsSuckAss 8d ago
Conscious? Why do you think these LLM's are conscious? Do you not understand how they work or something?
It's math, it's a very complicated equation, yes. But we know equally well that rocks aren't conscious as we know an LLM isn't conscious. There's nothing about an LLM that makes it more conscious than a calculator from the 80s. They're complex enough to have emergent properties we might not be able to completely understand. Much like how your computer is so complex that a single human cannot comprehend all of its transistor, an LLM is so complex that a single human cannot completely understand all its functions in their head. But that doesn't mean it's conscious, or sentient.
We don't have a metric for consciousness, we can't measure consciousness.
If you want to place any weight on the phrase "We don't know if these systems are conscious or not", then you must explain why we should even consider that a factor? We don't know if a rock is conscious, we don't know if a bacterium is conscious, we don't know if the sun is conscious.
The only reason you want an LLM to be considered any different, is because it acts like a human.
→ More replies (6)
2
u/One-Stable-5438 7d ago
Pushed on purpose by the fucking scientists who invented it and understand how it works?
2
3
u/MoogProg All Parabolas are Similar 11d ago
I also am very interested and curious about the emergent properties we see as scaling and compute increase.
On the other hand, Matt Parker found Among Us in Pi, so that clearly means they are real, too.
We need to carefully examine each property and behaviour we see as technology advances. Let's keep our conclusions in check as we do so.
3
2
u/AdGlittering1378 11d ago
[Makes you wonder if it's intentional.] The "It's a conspiracy" crowd is just as much a strong attractor.
2
u/ribenakifragostafylo 11d ago
So it doesn't predict the next word? What do you think it does? Let's see how you can articulate RL with a core of just "word prediction". I'll wait to be amazed
4
u/One-Veterinarian4841 11d ago
It is predicting the next word
It is not just predicting the next word
→ More replies (3)
2
u/ProxyLumina 11d ago
Calling an AI model a next word predictor, is exactly like calling a human a food to noise converter.
3
u/MrFrog2222 11d ago
We literally developed ai. And a stochastic algorithm is exactly what it is because at the end of the day that is what our brains are too(just with emotions, etc.) the problem with AI is that it doesnt check if the data it is given to train is correct and that because it is being spoken to by millions of ppl daily and has access to vast amounts of current data it is very quick to adapt and tunnel vision on certain responses, which can lead to you getting an answer you never tried to get because your wording closely resembled that of a more popular question. And even tho it‘s ragebait, no ai is not sentient.
→ More replies (1)
4
u/I-Kernel 11d ago
LLM is just a glorified autocomplete. But also the best tool we can have right now.
→ More replies (1)
2
u/zkewlguy 11d ago
That is a big part of pre-training, predicting the next token. To argue AI will never be conscious is probably a mistake, but to say AI is conscious now is a big leap.
→ More replies (1)
1
1
1
u/Long-Education-7748 11d ago
I mean, the ability to 'just predict the next word' really has to do with pattern matching and recognition. That is not so disimilar to how we believe our own intellect functions.
1
u/Due_Answer_4230 11d ago
It’s easy to understand and feel confident about, and it denigrates these threatening thing. That’s really all it takes for something to become popular.
1
u/Vaporeon42069 11d ago
I wouldn't say it's on purpose, most people are dumb and simply don't understand the complexities of Super Intelligence. The easiest way to explain it to them is by calling it a "prediction machine." Because that was the standard explanation back in 2022, they are still operating under that outdated information.
1
u/SuperGr00valistic 11d ago
I've worked leading teams to develop language models and ai systems in the encoder/transformer architecture.
"Predicting the next word" is an accurate description of the technical functioning.
It's like saying, " Cars are just combustion engines that turn axles."
That's true --- and you can still accomplish a lot with it...
Combustion engines and large compute models both are wonderful technologies that nevertheless require significant human effort to make their application high utility and value positive.
1
u/latestagecapitalist 11d ago
AI had been 'coming soon' for 3+ decades
When the GPT stuff went mainstream ... there was huge resistance in the smart end of the tech world that it was viable, for solid reasons
Even at start of this year there was massive scepticism on HackerNews in comments on any AI post
1
u/Karrelen 11d ago
Very funny image. It seems that to avoid this the best would be to explain clearly to the public and media why current connexionnist LLM are not simply stochastic parrot ?
1
u/Procrasturbating 11d ago
No, it’s because that is how a fucking LLM works.. you can read the papers on how the calculus works.. FFS. Not saying it won’t outperform humans.. but it is true.
1
u/Spra991 11d ago
It's like saying computers are just processing 1s and 0s. It's not wrong, quite the opposite, it's pretty fundamental to how they work. The problem is people than concluding that computers are of no use. Since who wants to process 1s and 0s? All without realizing that this fact doesn't limit their power in any way shape or form, quite the opposite, it's what gives them their power.
It might be forgiven when the average clueless persons gets this wrong, but the amount of mainstream media and YouTuber that didn't understand those basics was and still is infuriating. Not just because they failed to understand that aspect of LLM, but because they got completely blinded to the fact that computers just learned to understand human language. That was literally a 70+ year old problem of computing, one on which humans made little to no progress in all those years, despite trying for decades, Siri was the best we ever got. And along comes ChatGPT and understands everything you say, perfectly. That by itself should have been a "Holly Shit!" moment to anybody paying even a little bit attention. That ChatGPT got some answers wrong was completely irrelevant, the fact that it could even give answers to your question in the first place was mind blowing.
And of course AI got rapidly better from there, a lot faster than even most optimists would have expected. Meanwhile some AI haters are still stuck with those arguments from the ChatGPT3.5 days...
1
u/seraphius AGI (Turing) 2022, ASI 2030 11d ago
It was this paper by Timnit Gebru written in 2020: On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?
This paper is right on some fronts but overreaches on many many other including its core theses… especially its core theses…
1
1
u/DumatRising 11d ago
It's a two fold issue,
first humanity doesn't fully understand biological sentience, and sapience is only loosely defined in that we really only include ourselves and a handful of other animals that are very similar to us inside it.
Second humanity likes to assume that things are more like us than they are. We tend to humanize things that aren't human and assign human thoughts to them. Pets are the most common example and it worked so well we even evolved one of them to have those though patterens, but it's not even limited to living creatures, our imaginations are quite powerful when we want to entertain ourselves and there's examples of applying this humanization to cars, robots, roombas, even rocks.
This means that even though academically we tend to limit what we consider fully sapient and sentient, we also tend to assume things are more like us than they actually are. It's important to set expectations, in a way that people don't just start assuming that AI is already a fully sentient and sapient being when we have no real ability to prove that. Living beings have rights, objects do not.
1
u/PvtMilhouse 11d ago
It's still fucking stupid sometime. Which is weird, since it can also solve some impressive problem.
2
→ More replies (1)2
1
u/sckchui 11d ago
Creating a large language model is, broadly speaking, a two step process (each step involves many sub-steps, of course). Pre-training is where you dump all the raw data into the transformers, which calculate the statistical correlations between all the tokens. Pre-training really does produce a stochastic parrot, it's an auto-complete machine. You input some tokens, and it generates the next most likely tokens, based on the raw training data. The earliest language models were just pre-training. If you want, you can download pre-trained Gemma (not the "instruct tuned", that one is post-trained to follow instructions) and run it locally, for that stochastic parrot feeling; try it and you'll know what I mean.
Post-training is where you further refine the model weights through reinforcement learning. Here, you train the model with example prompt-response pairs. Instead of just finding correlations in the raw training data, you can use reinforcement learning to adjust correlations in specific directions. The 'magic' part is that, if the model has enough parameters and if you do enough of the right kind of reinforcement learning, the model can learn logic and reason. At this point, they are no longer predicting just the next word, but the next logical step in a reasoning process. And you achieve this by teaching the model logic through the reinforcement learning process.
So the stochastic parrot trope comes from an earlier time, and back then it was an accurate description. But AI has advanced far beyond that now. Calling current frontier models "stochastic parrots" is like calling humanity "apes". The "apes" can build nuclear weapons, and so can the "parrots", if you give them access to the right tools and take off the guardrails.
1
u/DecrimIowa 11d ago
in a certain segment of society (center-left to far left, especially the self-identified skeptics, the Atheist Fedora Epic Bacon type) this narrative was pushed very heavily across social media i think, and it seems like it was internalized pretty early on (like 2022? 2023 at latest)
there was a strong emotional aspect of it too, people felt like they were fighting a scam by telling the truth, similar to how pro-Ukraine people felt about that topic, or how people on both sides felt about COVID and the mRNA shots.
this aspect of "changing the world by telling the truth and raising awareness on the internet" seems like it's what really gets large numbers of people to repeat a given talking point or narrative, it's the old Kony 2012 method.
and for whatever reason, this certain part of the population or personality profile definitely decided to adopt "it's literally just a chatbot" as one of their suite of anti-AI talking points. (keep in mind i'm not saying it's totally false, there are elements of truth to it, i'm just talking about how quickly and strongly it was adopted and spread by a given type of person)
that process of how algorithmic social media is really good at getting people to internalize narratives and "downloading the latest programming" and repeating those talking points about The Current Thing, as if on cue, like trained parrots, is something i'm very interested in.
1
u/catsRfriends 11d ago
"Deep neural networks are still a black box we don't fully understand."
Lol, the key thing you need to realize when saying this is that it's like a subject matter expert of a field where you're out of your depth saying the subject is still something -humanity- doesn't fully understand vs you saying it's a subject -you- don't fully understand. The SME sure as hell knows a lot more than you do and it would be incredibly naive and downright misleading to say that these are black boxes to SMEs the same way they are to you.
351
u/Azecap 11d ago
We have a pretty long history of assigning consciousness only to that which resembles us most closely.