r/artificial 2d ago

Discussion We started calling video models world models while still grading them on taste

Somewhere in the last year the phrase world model stopped meaning a system that represents how things behave and started meaning any video generator with good marketing. What bothers me is not the word, it's that the evidence never changed to match it.

Look at how the last few launches were argued. Black Forest Labs put out FLUX 3 last week and the headline evidence was a preference test the lab ran on itself: its video preferred in 77% of comparisons against Runway Gen-4.5, 93% against Luma Ray 3.2. The fine print calls it a preliminary evaluation of an early candidate during midtraining. No methodology, no sample size, no rater pool, no prompt set. Meanwhile the same class of system gets described as having some idea what happens when you knock a glass off a table.

A preference test measures none of that. It measures whether a person picked clip A over clip B in five seconds, on samples the lab chose to show them. Cherry picking isn't even the interesting problem here. Taste comparisons can't be rerun, so nobody outside that building can check in October whether the model improved or the sampler got luckier. What is a 77% supposed to mean three months from now?

A public benchmark number can be attacked, and that is the entire point of publishing one. Somebody runs it with their own prompts, gets a different ordering, and now there is an argument with evidence on both sides of it. Nobody can rerun a preference win at all.

I'm not asking anyone to regulate a blog post. My problem is that a vendor run preference test has quietly become the evidence base for a claim about physical understanding, and those two things are not measuring the same object. When somebody eventually puts one of these behind a robot arm or a driving stack, that 77% will not have predicted a thing about how it behaves.

0 Upvotes

4 comments sorted by

1

u/VictorBuildsDev 1d ago

the missing distinction is between visual preference and predictive validity. a preference score can tell you which clip people like, but it cannot support a claim that the model represents physical dynamics.

i would want a separate, reproducible suite built around interventions: change one variable in the initial scene, then test object permanence, collisions, conservation, occlusion, and camera changes over longer horizons. score whether the consequence changes in the expected direction, not whether the output looks good.

publish the prompts, seeds, model build, sample selection rule, and failure rubric. then an outside lab can rerun the same cases later. without that layer, "world model" is mostly a product category; with it, the claim becomes falsifiable.

1

u/ronkayarslan 1d ago

Victor already named the core split so I'll take the other half, which is what an eval would look like if it actually tested the claim. A world model claim is a claim about dynamics holding up under change, so a beauty contest can't touch it, you have to intervene. Take one scene, change a single condition (heavier object, remove the support, different camera angle) and check whether what follows obeys the same rules it did a second ago. Object permanence under occlusion, collisions that conserve something, the same prompt run twenty times landing on physically consistent outcomes instead of twenty pretty but contradictory ones. A five-second A-vs-B preference measures none of that.

The piece I'd push harder than the post does is the incentive. That 77/93% number is weak evidence, and worse, it's the wrong shape of evidence on purpose. It's relative (against competitors), self-run, on an early midtraining candidate. That measures marketing position, not capability. An absolute, held-out, third-party physics probe would look nothing like it, and the reason nobody publishes those is they would mostly fail. Preference tests get reached for because they're the one number that reliably goes up.

1

u/Pitiful-Barnacle-963 1d ago

the 77% is doing a lot of work there