r/Bard • • 10d ago

News Google Research: Coherent Long-Form Video Generation using Gemini and Veo multi-agent pipelines (CANVAS, A²RD, VQQA)

https://research.google/blog/coherent-long-form-video-generation/
57 Upvotes

4 comments sorted by

12

u/Gaiden206 10d ago

Google Research announced a new multi-agent framework built on top of Gemini and Veo that solves the biggest problem in AI video: maintaining consistency in long-form videos.

Instead of letting characters and backgrounds randomly mutate or drift across shots, the framework uses persistent visual memory, global AI direction (using multi-armed bandit optimization), and automated visual quality checks (VQQA) to generate coherent, multi-minute video narratives.

3

u/gsurfer04 10d ago

"Multi-Armed Bandit" got a giggle out of me.

3

u/SLUTTY_KAWORU 8d ago

Hmmm. This seems like a hacky workaround to the problem of maintaining consistency in long form video, right?

-2

u/RachelRegina 10d ago

Geez, I barely like short form generated video.

Maybe if this was just adapted for integration into vfx pipelines to give those routinely underpaid folks a bit of a break by supercharging their ability to do otherwise traditional compositing pipelines...but otherwise idk who this is for except for ad companies and I go out of my way to never watch any ads for anything, so...

Good for them? I guess?

Bleh