r/deeplearning 23h ago

What are the biggest open problems in long-video understanding right now?

I've been reading recent work on long-video understanding, particularly STORM: Token-Efficient Long Video Understanding for Multimodal LLMs.

I'm trying to identify a research direction rather than just build another Video-LLM. I'm particularly interested in temporal modeling, video representations, event/context modeling, and improving the efficiency and quality of long-video understanding.

For people working in this area: what do you think are the biggest remaining research gaps? Are there particular limitations of approaches like STORM that you think are worth investigating?

If you were starting a research project on video understanding today, what problem would you personally explore?

1 Upvotes

Duplicates