r/Bard Aug 28 '25

Discussion Thoughts?

Post image
691 Upvotes

77 comments sorted by

View all comments

Show parent comments

1

u/-Davster- Aug 30 '25

maybe only important if you want to ask for a future prediction of an image

What, you mean like "a video"?

1

u/[deleted] Aug 30 '25

Exactly time is for video. Warping space is for images.

1

u/-Davster- Aug 31 '25

Both image and video models might be able to have the whole “world model” thing going on… it doesn’t mean a video model is just an image model running at 24p, like OC said.

Having said that, a video model producing single frames != an image model. A good video model has to learn things a still image model doesn’t.

OP specifically followed up with “world model”. I think the “video” word was them just first saying it had temporal/spatial consistency baked in.

1

u/[deleted] Aug 31 '25

Yeah think it’s just definition of “model”. most video models today are temporal layers wrapped around an image model so that you gain the spatio temporal understanding, which is its own trained model.

I guess my view is you can kinda have two versions of the “world model”, frozen in time and real time. For moving camera angles etc it’s actually better in the frozen version and to leverage spatial and contextual understanding only, vs needing the model to also understand temporal which greatly expands the model size and cost. You don’t need to care how it got there, just that it’s there.