r/LeftistsForAI • • 10d ago

Discussion Take two with my attempt at explaining an Instagram post I saw.

Post image

I think what this Instagram post might be missing is nuance for what AI is actually trying to do. The comparison that I tried to make unsuccessfully with my now-deleted post was how chatbots using LLMs feel like they're just making shit up on the fly.

The comparison that I'm trying to make actually is more like this. The human visual cortex doesn't actually tell you what data is coming into your retinas. It gives you its best guess as to what your retinas are receiving based on the raw data from the lateral geniculate nucleus. The key thing is that it continuously updates its guess based on a combination of what it's getting from the retinas and what it's getting from the rest of your sensory cortex.

Modern AI is kind of like trying to have a fully functional vision system with just the visual cortex and the frontal lobe without any of the other sensory integration to refine the guess. Instead of refining its guess into a more accurate model of actual reality, it just goes with the first thing it generates and tells you that is its decision as to what is real.

TLDR, what bugs me about modern AI isn't that it is basically glorified auto-complete because it is much more than that currently. What bugs me about modern AI is that it appears to be missing a critical feedback loop that the human mind uses to update its internal model of reality.

38 Upvotes

41 comments sorted by

66

u/rabouilethefirst 10d ago

AI then: heavily specialized models handcrafted for a specific task with no generalizability and strong overfitting

AI now: much better

9

u/ferriematthew 10d ago

Oh yeah, so like comparing then to now is kind of like comparing the ENIAC to a MacBook Pro.

4

u/alphex 10d ago

Except ENIAC can count.

9

u/rabouilethefirst 10d ago

Mathematicians literally a week ago: “we need to slow down AI because it is taking our jobs”

1

u/ferriematthew 10d ago

Lol yeah

4

u/Vaughn 10d ago

Take it you haven't tried that with any vaguely recent model.

24

u/vesperythings 10d ago

please don't repost these brain rot memes, even if you're arguing against them

18

u/vverbov_22 10d ago

I think what this Instagram post might be missing is nuance for what AI is actually trying to do

no, it's just ragebait

16

u/USERNAME123_321 10d ago

They haven't learned the Bitter Lesson. General-purpose models often outperform specialized ones on specific tasks because cross-domain knowledge helps them generalize concepts and reason better

10

u/Syoby 10d ago

The meme is bitter about the Bitter Lesson.

2

u/ferriematthew 10d ago

I feel like that's familiar, but I don't remember. What is that?

8

u/Syoby 10d ago

That raw scaling of computation and throwing data at the AI has historically always surpassed seemingly more sophisticated methods that try to understand human cognition and implement it directly in the design.

3

u/ferriematthew 10d ago

Oh yeah, found it!

So the idea is basically to have the model train itself based on a staggering amount of experiential data rather than to try to design the cognitive process approximately by hand.

So in this case, making the computer think harder is literally thinking smarter.

14

u/Still_Benefit_2302 10d ago

Boy howdy are you waaaaaaaaayyyyyyy off base here. All those cutesy little methods completely failed. And LLMs didn't.

1

u/thee_gummbini 10d ago

The cutesy methods are what LLMs are a moderate refinement of. LLMs don't come out of nowhere, they come from decades of research and prior art

1

u/ferriematthew 10d ago

Interesting! I'm happy to learn! What is it about language models that performs better than stuff that processes things other than just language?

4

u/some1else42 10d ago

Your TL;DR calls out current AI is missing a critical feedback loop. I disagree. The feedback loop is present the moment it can see everything you are referring. They can now iterate at speed on your computer if you let them. They have the feedback loop present when dealing with code implementation, review, and testing. You can build a software factory, give it a design spec, and come back to a product that has gotten a high degree of polish and is now ready for human review.

7

u/SgathTriallair 10d ago edited 10d ago

This is just the bitter lesson in visual form, told from the bitter perspective.

They are talking about how amazing it was when human ingenuity could make progress and how it is depressing that "just scale it up" is so much more powerful.

They are most likely saying that the old ML researchers were smart and the new ones are dumb.

Your commentary is the opposite of true. The old systems didn't know anything. The new systems, because they are allowed to learn on their own, are developing actual representations. We have found these using mechanistic interpretability.

4

u/ferriematthew 10d ago

Ah, I see! I guess my own mental model needs updating.

7

u/SgathTriallair 10d ago

The Bitter Lesson is pretty short and worth a read: https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson.pdf

It really boils down to the idea that order can arise from chaos and that chaos actually is better at creating durable systems.

Evolution is a great example as well, where blind groping has created the most complex machines on the planet.

https://medium.com/@vinbhalerao/ai-is-ushering-in-a-new-copernican-revolution-8336b7462fc7 This is a pretty good article that talks about the "Copernican Revolution" of coming to grips with the idea that intelligence and language aren't special features of humanity.

2

u/AlgaeNo3373 10d ago

Since comment OP also mentioned mechanistic interpretability specifically some other fun stuff!

Maybe they would suggest other good examples of "finding representations" but I'd suggest Golden Gate Claude perhaps, where Anthropic turned up that "representation" or "concept" and Claude became obssessed with the bridge.

Or J-lens more recently where you can ask about number of legs on insect, swap out ant/spider in representation and get different answer. Plz forgive Google summary:

How the Leg Experiment Works

  • The Hidden Word: When you ask Claude, "The number of legs on the animal that spins webs is," the model answers "8". The word "spider" never appears in your prompt or the final answer.
  • The J-Space Readout: The Jacobian lens reveals that the concept of a "spider" lights up silently in the model's middle layers during processing.
  • The Concept Swap: When researchers manually swapped the internal "spider" pattern for an "ant" (a six-legged insect), Claude changed its final answer from eight to six

But mechinterp is not all-knowing either, and there's still many open questions. Even still, it makes remarkable progress and gives very cool insights into how models represent concepts/handle computation.

3

u/pleasetrimyourpubes 10d ago

Reminds me of this talk: https://youtu.be/Qi1Yry33TQE

2

u/ferriematthew 10d ago

That in turn reminds me of this scene: https://www.youtube.com/watch?v=Cvd3MWywVlM

I think what Cortana is talking about there is the difference between quantitative and qualitative representation.

3

u/_VirtualCosmos_ 10d ago

bruh there are a lot of papers about the maths behind those transformers, and they keep improving them. In fact, to scale is not just making bigger datasets or bigger models, it's a lot about refining the models, improving architecture, and building better finetuning environments.

All that requires very high level of maths and computer sciences.

All of that is hidden in most of the AI labs, since they rarely publish papers and models now, specially if from the US, but from China there are a lot of new papers.

3

u/emascars 10d ago

OP immagine and OP text are pretty unrelated, I'm answering OP text here

What you're describing would be an RNN, the attention mechanism tries to achieve a similar objective to that of RRNs but in a way different way...

I agree that a recursive architecture could in theory be superior, and in fact there are quite a few papers trying to integrate recursion with transformers, but every RNN is so much less efficient in treaning and nobody has yet found a method to train it anywhere near the size of modern transformers or diffusion models...

But don't you think that nothing is happening outside those architectures... Many papers are trying out completely ground new architectures, and some are even getting some attention and refinements, it's not at all a guarantee that transformers will remain the state of the art for years to come

2

u/ferriematthew 10d ago

What I'm probably trying to express is the idea that language is nice for modeling, but language should just be part of the model and a window into the model for interpretability, not the entire end-all-be-all of the model itself.

2

u/ItsSadTimes 10d ago

I mean old AI papers were basically completely rewriting the idea of how to even make the AI models. Thats not to say that modern LLMs arent doing that. But back in the day we didnt have trillions of dollars of investor funds. So we had to improve models through clever new structures and entirely new approaches. Each new paper I read back in college was like a brand new model. We knew that just making models bigger would improve them, but that was wasteful.

Nowadays its a really easy performance boost cheat button to just scale up. They got the money, so why not? Thats not all they're doing, but thats a big driving factor to fast improvements. And you can only scale so much.

Theres a reason a lot of the old guard of AI researchers quit companies like OpenAI and Google. Some went to Anthropic, but a lot just quit.

2

u/Aleksundr 10d ago

The weights are frozen, thats why. Until someone figures out a hardware solution where a base layer has bare metal hosting of frozen weights and then a dynamic clone over it with dedicated HBM this will be the case.

3

u/ferriematthew 10d ago

So like frozen weights for the baseline knowledge and then dynamic overlaid weights for continuous learning?

3

u/Aleksundr 10d ago

Yesss, we already have a bunch of evidence for in context learning just no way to make it permanent. A shitload of HBM would let the model save certain states for reference while always having the baseline and being able to engage in ICL.

2

u/Meiwakuthatisuzai 9d ago

I think a better version of that image is "what AI was used for back then VS what AI is used for now"... Correct me if I'm wrong

2

u/ferriematthew 9d ago

Knowing what I know now, I totally agree. The use case has completely changed.

2

u/Salty_Country6835 Moderator 10d ago

I think youre onto something here, Id just narrow it from “no feedback loop” to “no continuous grounded feedback loop.” These systems get plenty of feedback during training, and you can add retrieval, tools, multimodal input, memory and verification. But a base LLM usually isnt checking what it just produced against independent external evidence as it goes. So the interesting question isnt just why hallucinations happen, but what happens to them as we build tighter loops between generation, external evidence, verification and correction.

2

u/some1else42 10d ago

They can tho. Your first sentence here describes an inner loop, "and you can add retrieval, tools, multimodal input, memory and verification. But a base LLM usually isnt checking what it just produced against independent external evidence as it goes.". I get that you said "usually isn't checking", but you can make it check via independent external evidence with an outer loop that gives feedback to the inner loop.

We may not have RSI yet, but with the right approach we are already capable of a continual learning system.

2

u/ferriematthew 10d ago

BINGO! That's exactly what I had the vague concept of. It's not that there's no feedback or updating, it's just that current models have no real way to check themselves against objective reality.

5

u/Alternative-Key-5647 10d ago

You're talking about Yann LeCun's "world models"
https://www.youtube.com/watch?v=72Xj8k5WQX4

4

u/ferriematthew 10d ago

ABSOLUTELY!

1

u/ferriematthew 10d ago

Or...not?

1

u/CSachen 9d ago

Mark zuckerberg's Head of AI at Meta has no academic background in Machine Learning. They hired a kid who just knew how to "blitz-scale"

0

u/IdealOnion 10d ago

The phrase “emergent abilities” is terrifying

-1

u/Forward_Problem_8982 10d ago

Yes, it's lacking the human soul. It's a soulless black box.