Publication Type

Journal Article

Version

acceptedVersion

Publication Date

2-2026

Abstract

In the realm of video dialog response generation, capturing both the essence of video content and the temporal nuances of conversation history is crucial. While some approaches rely on large-scale pretrained visual-language models, often neglecting temporal dynamics, others emphasize spatial-temporal relationships within videos but demand intricate object trajectory pre-extractions and overlook dialog temporal dynamics. This paper introduces the Dual Temporal Grounding-enhanced Video Dialog model (DTGVD), designed to bridge the gap between these two approaches. DTGVD uniquely integrates the strengths of both by emphasizing dual temporal relationships. It achieves this by predicting dialog turn-specific temporal regions, selectively filtering video content, and grounding responses in both video and dialog contexts.A key innovation of DTGVD is its advanced handling of chronological interplay within dialogs. By effectively capturing and leveraging dependencies between dialog turns, it enables a more nuanced understanding of conversational dynamics. To further align video and dialog temporal dynamics, we introduce a list-wise contrastive learning strategy. In this framework, accurately grounded turn-clip pairings are treated as positive samples, while less precise pairings serve as negative samples. This refined classification is then seamlessly integrated into our end-to-end response generation mechanism.Evaluations using AVSD@DSTC-7 and AVSD@DSTC-8 datasets underscore the superiority of our methodology.

Keywords

Temporal Grounding, Video dialog, Vision-language models

Discipline

Artificial Intelligence and Robotics

Research Areas

Intelligent Systems and Optimization

Areas of Excellence

Digital transformation

Publication

IEEE Transactions on Multimedia

Volume

28

First Page

7151

Last Page

7162

ISSN

1520-9210

Identifier

10.1109/TMM.2026.3668656

Publisher

Institute of Electrical and Electronics Engineers

Additional URL

https://doi.org/10.1109/TMM.2026.3668656

Share

COinS