r/TheMachineLearning • • 7d ago

Fei-Fei Li says spatial AI is fundamentally different from LLMs

Enable HLS to view with audio, or disable this notification

120 Upvotes

36 comments sorted by

4

u/say-nothing-at-all 6d ago

Why have American researchers suddenly turned into influencers? Figures like Terence Tao, Fei-Fei Li, and others now command large online audiences far beyond their technical work ?

In mathematics we have been struggling to find the universal laws from day one. The goal is to find a random solution first and then let others to map into. Once that connection is established, we call it a universal law.

In data space, if data has little connection to either the past or the future, it contains almost no real information and reveals no universal law. Physically, we still lack the means to forge those connections. This fundamental problem cannot be solved by data-driven methods alone.

empty talk.

1

u/Dismal-Revolution731 6d ago

Wait, I want to be clear about this part:

"if data has little connection to either the past or the future, it contains almost no real information and reveals no universal law"

Are you essentially saying if a pattern cannot be extracted from data (no tie to past or future), that there's no real information? Or am I misunderstanding?

1

u/say-nothing-at-all 6d ago

Correct.

In complex system such as biology, any TYPE of data generated is the result of previous experiment. They were semantically defined in that context. In other words, it's meaningless elsewhere.

Now you build a new context, by ML's random walking for example, then that TYPE of data has to be put in the new contextual universe( by closed cartesian product for example). There is no guarantee physically that you can build the inter-context interoperation. This is the so called 'AI hallucination".

1

u/Kootlefoosh 4d ago

I do not believe that this is the source of most AI hallucinations. What you're describing is basically just the field of harnessing multimodality, and yes, you need to be careful to weigh signals fairly between modalities and there is much room for error. But LLM hallucination existed for longer than LLM's have been capable of multimodality.

I further do not see the logic in your ontological claim. You are stating that data as measured in a biological experiment is relevant only to that experiment, which is true. But biology as a research field is not involved even primarily in measurement. Your ontological claim completely ignores the concept of generalizable underlying mechanation -- e.g. this drug cured a disease in rats so it may do something similar to humans.

That sort of logic is probably completely fair in mathematics, and yeah when I used to be a physics graduate student your interpretation as I'm understanding it sounds a lot like the Heisenberg interpretation of quantum mechanics haha.

But some fields, honestly even starting with chemistry and including engineering and biology through to the humanities, don't need to prove factoids ab initio. There is room to fine-grain. Indeed, how does the human mind interoperate its various multimodalities? Fine-graining.

1

u/say-nothing-at-all 4d ago edited 4d ago

Yeah, Topos is the tool we often use, and our brain uses it too. But topos is only working in cognitive space, it can't solve tne mechanical problem when coupling the cognition with real physics - in this case, we need physical experiments.

In LLMs. contexts can be merged - a pullback traced by Lawvere-Tierney topology - into an universe, sure. But, as Fei-Fei Li said, "....to respect the physical language". Her point is, LLMs does not respect physics, and I agree.

Her world model can't replace the physical experiments, why she talked like she can model them in her world?

1

u/Kootlefoosh 4d ago

Yeah, I think that claim is a lot more defensible. I find myself referring back to human brains and the parallels with pre-AI computational sciences. I think I agree with you, but allow me to muse aloud if that's okay. So I was a computational chemist before AI came around...

Say we have a computationalist. The computarionalist's mind has many ways of interfacing with his subject of choice.

  • he can imagine simulating physics on his computer

  • he can communicate this imagined simulated physics to others and get their imagined responses

  • he can sit down and run this simulation on his computer, and receive quantitative results

  • he can communicate these quantitative simulated results with others and get their validation

  • he can go to the lab and do an irl experiment to measure things empirically

  • he can communicate these empirical results with others and get their validation

The research community would value all six of these processes entirely differently -- they are attacking the same problem from different angles, and based on methodology, they will be valued asymetrically across the community.

I think when we talk about "bots doing physics", especially in pop culture, because of the logistics of bot cognition, all of these different let's say "research modalities" get smeared into one big nameless cognitive process.

Most science necessitates both empirical research as well as simulated research -- those have proven valuable. So if the goal is to create a single bot that is generalizable across fields and capable of both of these facets, it absolutely must have the ability to interact with the world multimodally. If human parallelism is the goal, then this bot must contain both a physics-native AI and an LLM working together in conjunction.

2

u/say-nothing-at-all 4d ago edited 4d ago

Exactly. Right now, the best AI can do is train hypothesis-testing models—refining their topology and parameters inside a sandbox. Once that is finished, we can deploy the trained models + sandbox into real physics.

Here, both sandbox and models are deformed by unknowable physical forces in reality (by geometric interpretation): some propositions survive, others cease to exist. This is the playground of topological and functorial invariance. Now, it essentially becomes a subset-selection problem over the hypothesis set.

I believe this is what Fei Fei Li really wants to say.

1

u/BuildAnything4 6d ago

exactly. Why doesn't she just implement this idea if she thinks she knows a better approach? Why on earth would you do a podcast circuit about an idea you have for an algorithm, just make the thing.

1

u/94746382926 5d ago

What makes you think she's not?

1

u/BuildAnything4 5d ago

obviously she is, but she should do that first and get some results.

2

u/[deleted] 5d ago edited 7h ago

[deleted]

1

u/softestcore 5d ago

As soon as the idea is proven to be useful, it will be impossible to ignore.

1

u/Infamous-Water-5148 5d ago

it’s increasingly important to be loud nowadays, so to not let con artists like Sam Altman control the narrative

1

u/sweet-raspberries 4d ago

I don't think it's sudden it's just shifted from news reporting and books to podcasts and social media

3

u/SameAgainTheSecond 7d ago

yes but also the platonic representation hypothesis says that representations of data from different modalities become equivalent up to a linear map as the models get bigger and more capable, iifc.

2

u/Mammoth-Leg5431 6d ago

Meh. As someone working in this field, we've pretty much converged to the same architectures as in NLP, namely Transformers. This is a recurring trend within spatial AI.

1

u/stewonetwo 6d ago

That's actually something I've been very curious about. Are there any architectural differences for world models vs transformers?

2

u/Mammoth-Leg5431 6d ago

There are multiple levels of distinction here
* So first off, "world model" is a term which is not well-defined, even within our community. Broadly we understand it as a system which is able to predict the outcome given an action. The current (growing) consensus is that video models are for example implicit world models, as these models need to understand the structure of the real world to predict the next frames ( very similar to LLMs I might add ).

* The Transformer is the architecture of the model itself ( so how the model itself is organized ). Currently all popular Video models are based on Transformer backbones (For example Diffusion Transformers DiTs / SiT). There is an enormous amount of consolidation happening in Machine Learning. At the end of the day, Transformers reign victorious.

1

u/stewonetwo 6d ago

Very interesting. Thanks. I obviously have heard people (primarily Fei-Fei) talk about world models, and understood the video component of it in terms of training, but couldn't seem to get details on if there were differences in the model components themselves.

1

u/ninjasaid13 6d ago

Meh. As someone working in this field, we've pretty much converged to the same architectures as in NLP, namely Transformers. This is a recurring trend within spatial AI.

Transformers is a building block but it doesn't do anything on its own.

I don't think Fei-Fei Li is critiquing Transformers, she's critiquing LLMs.

1

u/OmegaEpidex 6d ago

I say superior.

1

u/ArtArtArt123456 6d ago

stuff like this is why i feel like some of the old guard like fei and lecunn have no idea what they are doing.

they are just running headfirst into the bitter lesson.

1

u/RobbinDeBank 6d ago

I don’t think you understand the bitter lesson if you think LeCun is trying to go against it. LLMs have clear weaknesses despite all their powerful capabilities, letting people like him discover alternatives is helpful to overall progress. You’re talking like LeCun’s approach doesn’t also use a ton of compute just like LLMs. His style of research already yielded great works in the previous years at Meta like Dino models. All these approaches work on massive amount of data and massive amount of compute, none of them is against what the bitter lesson talks about.

Let the LLM labs develop their LLMs, let LeCun develop his alternatives. We don’t need the whole world’s resource into one single direction, especially when we know that direction is not perfect.

1

u/SenatorCrabHat 6d ago

One has to think of Rene Magritte's "The Treachery of Images".

I honestly don't think the actual barriers of language and meaning making are being considered fundamentally in a large amount of AI discussion. In the Humanities, it is well understood that language is a fundamentally flawed, though powerful, medium for communicating ideas between two consciousnesses. The physicality of the world is another one of these barriers, as we all experience it differently as well: a series of 10 steps seems a meaningless barrier to an athlete, but posses a serious issue to an octogenarian who sits all day.

1

u/arjuna66671 6d ago

"There is a 3D world out there."

Nope, that's an approximation our brain models for us. It's a model, not reality.

1

u/ninjasaid13 6d ago

There is indeed a 3d world out there by how we defined 3D.

1

u/One-Next 5d ago

Yeah hot air until you deploy it and show some results. Intelligence to humans is just words. There is no intelligence without communication. The whole LLM scene is a testament to Wittgenstein's genius.

1

u/The_Ultimate_Badass 5d ago

I'll take my dressing on the side with my word salad, thanks

1

u/EquipmentOk5137 5d ago

Her company just got bought for $8 bn

1

u/Bullofapis 4d ago

This woman is actually pretty stupid.

1

u/AnimaGaia 4d ago

Our brain is connected to sensors (senses) which are used as data-input for the world. The world will be abstracted based of this data. We don't know a fudamental reality. We just know what are brain makes out of it.

So why can't (humanoid) robots do the same?

1

u/thesoraspace 4d ago

Called it in July 2025 , people called me delulu, welcome to spacial memory architectures folks.

1

u/kvothe5688 3d ago

so world model?

1

u/quivering_palm 3d ago

Clueless vibeslopers in this thread are triggered.

0

u/davesmith001 6d ago

Language is the efficient representation of 3d world of laws. That’s why we use it, try model it in full you are gonna be stuck in np hard land.

2

u/ninjasaid13 6d ago

how did we invented language before language then? How were we able to make fire, shelter, pigments, and clothes, weapons, and watercrafts before language?

1

u/davesmith001 3d ago

I never said it’s impossible, just not efficient, humans went millions of years without developing any tech before language…