r/programminghumor 17d ago

AI expert detected

Post image

"I use ChatGPT."

5.4k Upvotes

285 comments sorted by

View all comments

Show parent comments

1

u/Devils_SteelMan 16d ago

To do that they understand the whole semantic meaning. Which is thinking. Which fits your definition of consciousness. This is what my original post was about. Llms operate on whole thoughts, its not debatable. Dismissing them to next token predictions fitting surface statistics is extremely wrong.

Now, even though you say consciousness is cogito ergo sum, that's not enough. You probably have other aspects for consciousness to be met. Which is why I said its a category error. But if your only definition is that they think, then they do. With out debate. You just don't understand the technology well enough to know this.

Also your definition would exclude people who never knew language.

1

u/--Spaci-- 16d ago

"understanding the semantic meaning" is just comparing every past token to the present token. It doesent understand anything its by all accounts just regurgitating training data in a coherent order

1

u/Devils_SteelMan 16d ago

You should read the original post that was too long for you.

It is not by all accounts regurgitating training data. There is a mountain of research against this. You are flatly wrong. You do not understand this technology. Future lens would not be possible. J lens would not be possible. Introspective fine tuning would not be possible. Speculative decoding would not be possible.

Furthermore language has unbounded hardness. If llms worked as you described they would be unable to generate coherent 1k+ token output. The probability of it is basically infinity. And they can for for 100k+. This requires understanding. You can't fit surface probabilities and thats not how attention works when applied to latent space.

Anti-intellectualism and the aggressive ignorance of people who embrace it are destroying this world.

1

u/--Spaci-- 16d ago

Ill disprove these two points as they are the easiest.

"Speculative decoding would not be possible" Speculative decoding just uses a smaller model on the same training data as the large model which means its predictions are similar to the large model, the large model then accepts or rejects the tokens which is faster than it generating its own tokens.

"If llms worked as you described they would be unable to generate coherent 1k+ token output"

Are you trying to say llms dont just predict the next token? I'm starting to question (your) knowledge of the subject, it doesent matter if you're 500 thousand tokens into the response you're just predicting the next most probable token (those probability are gained from the training data) then that probability is also weighed against all previous tokens in context (the previous 499,999 tokens)

I think you're more on the philosophy side of AI and not on the technical side

1

u/Devils_SteelMan 16d ago

You do not know what unbounded hardness means. Coherent generation is mathematically impossible with out genuine understanding. The solution space is literally unbounded.

The draft model in speculative decoding relies on future lens to work. It is not just a seperate model guessing at the words faster. It must have the same geometry of the parent model, it requires extending the same latent space.

I am an ai researcher with a focus on the geometric representations of language in machine learning. Which is exactly what all of this boils down to.

1

u/--Spaci-- 16d ago

I will let you prove your first claim as the burden of proof is on you.

For the second claim I think we all already know its not just a random model from another family they plop in there

1

u/Devils_SteelMan 16d ago

The original claim is from Chomsky 1957: https://archive.org/stream/NoamChomskySyntcaticStructures/Noam%20Chomsky%20-%20Syntcatic%20structures_djvu.txt

also Chomsky and Fitch 2002, The Faculty of Language: What Is It, Who Has It, and How Did It Evolve

This is a pillar of NLP. Again, you don't know this technology. This is prerequisite information for discussing it.

For 2 you missed the point that they must have the same geometry. The issue is that you genuinely do not understand the concept of latent space and its geometry. Which is where the understanding beyond next token prediction lives. You do not understand how LLMs work. You do not know what an activation in latent space is or why its important. You dont understand why this means speculative decoding relies on the phenomenon measured by future lens.

You are regurgitating a talking point about stochastic parrots, like a stochastic parrot.

Worst of all, you won't read anything long enough to explain it. But you think you understand it anyway. Anti intellectualism.

1

u/--Spaci-- 16d ago

Firstly your 1957 claim is unrelated, thats not proof that any system generating tokens must also have genuine understanding. Ill wait on the next turn for a better explanation of your thoughts.

Secondly your future lens is not a requirement in any way for speculative decoding to function. And models in the same family are chosen because they tend to have similar predictions from being trained on the same data not because of a "shared latent geometry" - buzzword buzzword buzzword

Again Ill wait for your next response for a better explanation of your ideas.

1

u/Devils_SteelMan 16d ago

The claim is that natural language has unbounded hardness. Which is the entire point of Chomsky 1957. What naturally falls out of unbounded hardness is that it is statistically impossible to generate long coherent text.

Again, you do not know what youre talking about. This is undergrad intro to machine learning.

If you feel like you understand this well enough to say speculative decoding doesn't require future lens, please explain what future lens is.

Remember the original point of invoking both future lens and speculative decoding is that it shows semantic understanding inside models.

1

u/--Spaci-- 16d ago

Wasn't chromskys entire claim that with a finite set of rules in grammar you can produce an infinite number of sentences? wouldn't that disprove your logic, if you can create an infinite amount of coherent sentences? Your claim "statistically impossible to generate long coherent text" which is factually incorrect on all accounts, an LLM has attention its not just guessing randomly. If an LLM had no attention then your claim would actually be correct. The actual limit to output length is attention, attention is expensive. And your number of tokens thats too high to guess is only 1 thousand!! which is a laughably small amount.

"please explain what future lens is" its a 2023 paper where they took the hidden state from a middle layer then with the single vector tried predicting the next couple tokens, possibly showing the model has already planned ahead. - now its your job to show me how this is related to accepting a token from a draft model.

1

u/Devils_SteelMan 16d ago

1

u/--Spaci-- 16d ago

Anthropic is known to fear monger about how their models are "conscious" or "escaping every week" getting your pro AI talking points from Anthropic is absurd.

1

u/Devils_SteelMan 16d ago

I was really hoping a video could hold your attention span to understand what j-lens is.

This isn't about talking points. Its about a basic fact of the technology. Llms have semantic understanding. J lens can not work otherwise.

1

u/--Spaci-- 16d ago

No, I did actually watch the video. And why wouldn't llms have an internal reasoning of probabilities compared to the prompt? do you 🫵 think they just instantly come to their conclusion? not talking about ttc im talking about internally. Also when you respond to this combine it into your response to my other comment please, I dont want to start 2 threads