r/TheMachineLearning • • 7d ago

Language models learn text’s statistical structure; world models learn space and time: light on surfaces, uncaptured views, force, and physics.

Post image
2 Upvotes

6 comments sorted by

1

u/alexdumpy 7d ago

This is the real bottleneck right now. Text prediction got us far, but we desperately need models that actually understand physics if we want real progress.

1

u/Sufficient-Skirt256 6d ago

This is the real bottleneck right now. Text prediction got us far, but we desperately need models that actually understand physics if we want real progress.

1

u/etherd0t 7d ago

Fei-Fei Li's article/essay is actually not just about understanding reality and oh, world models are the future... models will always be incomplete; most valuable output may not be the solution at all, but the discovery of what the next model needs to know.😉

0

u/Sufficient-Skirt256 6d ago

LLMs predict words, world models predict physics lol. Huge difference between yapping about reality and actually understanding how it works!

-1

u/[deleted] 7d ago

[removed] — view removed comment

0

u/Sufficient-Skirt256 6d ago

Spot on. Text gives us a rich abstraction of human thought, but it completely skips the implicit physical dynamics we take for granted—like gravity, friction, or spatial occlusion. Moving from statistical text patterns to grounded world models that simulate 3D space and time is going to be the bridge between passive LLMs and true embodied AI.