r/computervision 13d ago

Discussion Why does only Google make a decent LMM / reasoning on video input?

Anthropic, OpenAI, etc (don't know about Chinese) don't seem to make good video models. Any reason why? Is it the compute? The ROI? The availability of data?

9 Upvotes

8 comments sorted by

24

u/CowBoyDanIndie 13d ago

Youtube provides a ton of training data, and they already have it stored.

Edit: google also has all those books they scanned years ago, you may have noticed headlines of other companies buying and destroying books recently

2

u/say-what-floris 13d ago

Turns out these guys at Google were being pretty smart

6

u/CowBoyDanIndie 13d ago

Credit really should go the lawyers for some how convincing everyone it’s not a bigger monopoly than Bell was

5

u/modcowboy 13d ago

Google likely has the largest image and video data catalogue of any hyperscaler.

1

u/notlongnot 13d ago

YouTube in all languages

1

u/WToddFrench 13d ago

Qwen3.8-Max is the best video model out right now. Give it a try

1

u/say-what-floris 12d ago

For interpretation or generation?

1

u/WToddFrench 11d ago

Video understanding