Publication Type
Conference Proceeding Article
Version
publishedVersion
Publication Date
7-2026
Abstract
Video Large Language Models (Vid-LLMs) have demonstrated remarkable performance in video understanding tasks, yet their robustness under conversational interaction remains largely underexplored. In this paper, we identify spatiotemporal sycophancy, a failure mode in which Vid-LLMs retract initially correct, visually grounded judgments and conform to misleading user feedback under negation-based gaslighting. Rather than merely changing their answers, the models often fabricate unsupported temporal or spatial explanations to justify incorrect revisions. To systematically investigate this phenomenon, we propose a negation-based gaslighting evaluation framework and introduce GasVideo-1000, a curated benchmark designed to probe spatiotemporal sycophancy with clear visual grounding and temporal reasoning requirements. We evaluate a broad range of state-of-the-art open-source and proprietary Vid-LLMs across diverse video understanding tasks. Extensive experiments reveal that vulnerability to negation-based gaslighting is pervasive and severe, even among models with strong baseline performance. While prompt-level grounding constraints can partially mitigate this behavior, they do not reliably prevent hallucinated justifications or belief reversal. Our results indicate that current Vid-LLMs lack robust mechanisms for maintaining grounded spatiotemporal beliefs under adversarial conversational feedback.
Discipline
Artificial Intelligence and Robotics | Information Security
Research Areas
Intelligent Systems and Optimization
Areas of Excellence
Digital transformation
Publication
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), San Diego, California, July 2-7
First Page
14836
Last Page
14852
Identifier
10.18653/v1/2026.findings-acl.729
Publisher
ACL
City or Country
San Diego, California
Citation
TANG, Ziyao; JIAO, Pengkun; ZHU, Bin; QI, Huiyan; CHEN, Jingjing; and JIANG, Yu-Gang.
Spatiotemporal sycophancy: Negation-based gaslighting in video large language models. (2026). Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), San Diego, California, July 2-7. 14836-14852.
Available at: https://ink.library.smu.edu.sg/sis_research/11172
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Additional URL
https://doi.org/10.18653/v1/2026.findings-acl.729