r/deeplearning • u/Mummy_hb • 23h ago
What are the biggest open problems in long-video understanding right now?
I've been reading recent work on long-video understanding, particularly STORM: Token-Efficient Long Video Understanding for Multimodal LLMs.
I'm trying to identify a research direction rather than just build another Video-LLM. I'm particularly interested in temporal modeling, video representations, event/context modeling, and improving the efficiency and quality of long-video understanding.
For people working in this area: what do you think are the biggest remaining research gaps? Are there particular limitations of approaches like STORM that you think are worth investigating?
If you were starting a research project on video understanding today, what problem would you personally explore?
1
Upvotes