arXiv cs.AIPaper
Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning
Solid multi-modal work on a specific task. If you're building video understanding pipelines and dense captions matter, this approach to grounding temporal boundaries might be better than fixed assumptions. For most teams, this is specialist material.