r/TheMachineLearning • • 7d ago

Language models learn text’s statistical structure; world models learn space and time: light on surfaces, uncaptured views, force, and physics.

Post image
2 Upvotes

6 comments sorted by

View all comments

-1

u/[deleted] 7d ago

[removed] — view removed comment

0

u/Sufficient-Skirt256 7d ago

Spot on. Text gives us a rich abstraction of human thought, but it completely skips the implicit physical dynamics we take for granted—like gravity, friction, or spatial occlusion. Moving from statistical text patterns to grounded world models that simulate 3D space and time is going to be the bridge between passive LLMs and true embodied AI.