Publication Type
Conference Proceeding Article
Version
publishedVersion
Publication Date
7-2026
Abstract
Video Corpus Moment Retrieval (VCMR) requires models to efficiently retrieve and precisely locate specific moments relevant to natural language queries within a massive, untrimmed video corpus. However, existing discriminative approaches typically rely on shallow visual-textual feature matching mechanisms, which often struggle to capture fine-grained semantic differences. To address this limitation, we propose Video-GAR, a novel framework that reframes the conventional retrieval task from superficial matching to generative understanding, positing that the capability for query reconstruction evidences deep semantic comprehension. Specifically, Video-GAR orchestrates three synergistic components: To overcome the computational efficiency bottleneck, we construct a Bi-Mamba backbone that leverages the linear complexity of state-space models for efficient global context modeling. Building on these representations, we introduce a generation-augmented fusion module, in which a training-only decoder acts as a semantic regularizer to implicitly calibrate cross-modal attention without increasing inference overhead. Finally, to ensure fine-grained precision, we propose a boundary-aware localization strategy that integrates boundary modeling with categorical supervision. Experiments on two benchmark datasets demonstrate that Video-GAR significantly improves retrieval and localization accuracy while maintaining outstanding inference speed.
Keywords
Generation-augmented retrieval, Multimodal retrieval, Video corpus moment retrieval, Video understanding
Discipline
Artificial Intelligence and Robotics | Graphics and Human Computer Interfaces
Research Areas
Intelligent Systems and Optimization
Areas of Excellence
Digital transformation
Publication
SIGIR '26: Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, Melbourne, Australia, July 20-24
First Page
857
Last Page
868
ISBN
9798400725999
Identifier
10.1145/3805712.3809742
Publisher
ACM
City or Country
New York
Citation
KUAI, Mingjin; XIAO, Qianyin; LI, Juncheng; PENG, Jin; LIAO, Lizi; and JI, Wei.
Generation-augmented video corpus moment retrieval. (2026). SIGIR '26: Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, Melbourne, Australia, July 20-24. 857-868.
Available at: https://ink.library.smu.edu.sg/sis_research/11318
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Additional URL
https://doi.org/10.1145/3805712.3809742
Included in
Artificial Intelligence and Robotics Commons, Graphics and Human Computer Interfaces Commons