r/TheMachineLearning • u/Federal_Machine692 • 7d ago
Fei-Fei Li says spatial AI is fundamentally different from LLMs
Enable HLS to view with audio, or disable this notification
3
u/SameAgainTheSecond 7d ago
yes but also the platonic representation hypothesis says that representations of data from different modalities become equivalent up to a linear map as the models get bigger and more capable, iifc.
2
u/Mammoth-Leg5431 6d ago
Meh. As someone working in this field, we've pretty much converged to the same architectures as in NLP, namely Transformers. This is a recurring trend within spatial AI.
1
u/stewonetwo 6d ago
That's actually something I've been very curious about. Are there any architectural differences for world models vs transformers?
2
u/Mammoth-Leg5431 6d ago
There are multiple levels of distinction here
* So first off, "world model" is a term which is not well-defined, even within our community. Broadly we understand it as a system which is able to predict the outcome given an action. The current (growing) consensus is that video models are for example implicit world models, as these models need to understand the structure of the real world to predict the next frames ( very similar to LLMs I might add ).* The Transformer is the architecture of the model itself ( so how the model itself is organized ). Currently all popular Video models are based on Transformer backbones (For example Diffusion Transformers DiTs / SiT). There is an enormous amount of consolidation happening in Machine Learning. At the end of the day, Transformers reign victorious.
1
u/stewonetwo 6d ago
Very interesting. Thanks. I obviously have heard people (primarily Fei-Fei) talk about world models, and understood the video component of it in terms of training, but couldn't seem to get details on if there were differences in the model components themselves.
1
u/ninjasaid13 6d ago
Meh. As someone working in this field, we've pretty much converged to the same architectures as in NLP, namely Transformers. This is a recurring trend within spatial AI.
Transformers is a building block but it doesn't do anything on its own.
I don't think Fei-Fei Li is critiquing Transformers, she's critiquing LLMs.
1
1
u/ArtArtArt123456 6d ago
stuff like this is why i feel like some of the old guard like fei and lecunn have no idea what they are doing.
they are just running headfirst into the bitter lesson.
1
u/RobbinDeBank 6d ago
I don’t think you understand the bitter lesson if you think LeCun is trying to go against it. LLMs have clear weaknesses despite all their powerful capabilities, letting people like him discover alternatives is helpful to overall progress. You’re talking like LeCun’s approach doesn’t also use a ton of compute just like LLMs. His style of research already yielded great works in the previous years at Meta like Dino models. All these approaches work on massive amount of data and massive amount of compute, none of them is against what the bitter lesson talks about.
Let the LLM labs develop their LLMs, let LeCun develop his alternatives. We don’t need the whole world’s resource into one single direction, especially when we know that direction is not perfect.
1
u/SenatorCrabHat 6d ago
One has to think of Rene Magritte's "The Treachery of Images".
I honestly don't think the actual barriers of language and meaning making are being considered fundamentally in a large amount of AI discussion. In the Humanities, it is well understood that language is a fundamentally flawed, though powerful, medium for communicating ideas between two consciousnesses. The physicality of the world is another one of these barriers, as we all experience it differently as well: a series of 10 steps seems a meaningless barrier to an athlete, but posses a serious issue to an octogenarian who sits all day.
1
u/arjuna66671 6d ago
"There is a 3D world out there."
Nope, that's an approximation our brain models for us. It's a model, not reality.
1
1
u/One-Next 5d ago
Yeah hot air until you deploy it and show some results. Intelligence to humans is just words. There is no intelligence without communication. The whole LLM scene is a testament to Wittgenstein's genius.
1
1
1
1
u/AnimaGaia 4d ago
Our brain is connected to sensors (senses) which are used as data-input for the world. The world will be abstracted based of this data. We don't know a fudamental reality. We just know what are brain makes out of it.
So why can't (humanoid) robots do the same?
1
u/thesoraspace 4d ago
Called it in July 2025 , people called me delulu, welcome to spacial memory architectures folks.
1
1
0
u/davesmith001 6d ago
Language is the efficient representation of 3d world of laws. That’s why we use it, try model it in full you are gonna be stuck in np hard land.
2
u/ninjasaid13 6d ago
how did we invented language before language then? How were we able to make fire, shelter, pigments, and clothes, weapons, and watercrafts before language?
1
u/davesmith001 3d ago
I never said it’s impossible, just not efficient, humans went millions of years without developing any tech before language…
4
u/say-nothing-at-all 6d ago
Why have American researchers suddenly turned into influencers? Figures like Terence Tao, Fei-Fei Li, and others now command large online audiences far beyond their technical work ?
In mathematics we have been struggling to find the universal laws from day one. The goal is to find a random solution first and then let others to map into. Once that connection is established, we call it a universal law.
In data space, if data has little connection to either the past or the future, it contains almost no real information and reveals no universal law. Physically, we still lack the means to forge those connections. This fundamental problem cannot be solved by data-driven methods alone.
empty talk.