Publication Type
Conference Proceeding Article
Version
publishedVersion
Publication Date
7-2026
Abstract
Text-to-video (T2V) generation models have made rapid progress in producing visually high-quality and temporally coherent videos. However, existing benchmarks primarily focus on perceptual quality, text–video alignment, or physical plausibility, leaving a critical aspect of action understanding largely unexplored: object state change (OSC) explicitly specified in the text prompt. OSC refers to the transformation of an object’s state induced by an action, such as peeling a potato or slicing a lemon. In this paper, we introduce OSCBench, a benchmark specifically designed to assess OSC performance in T2V models. OSCBench is constructed from instructional cooking data and systematically organizes action–object interactions into regular, novel, and compositional scenarios to probe both in-distribution performance and generalization. We evaluate six representative open-source and proprietary T2V models using both human user study and multimodal large language model (MLLM)–based automatic evaluation. Our results show that, despite strong performance on semantic and scene alignment, current T2V models consistently struggle with accurate and temporally consistent object state changes, especially in novel and compositional settings. These findings position OSC as a key bottleneck in text-to-video generation and establish OSCBench as a diagnostic benchmark for advancing state-aware video generation models. Project page: https://hanxjing.github.io/OSCBench.
Discipline
Artificial Intelligence and Robotics | Graphics and Human Computer Interfaces
Research Areas
Intelligent Systems and Optimization
Areas of Excellence
Digital transformation
Publication
Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), San Diego, California, July 2-7
First Page
30867
Last Page
30884
Identifier
10.18653/v1/2026.acl-long.1425
Publisher
ACL
City or Country
San Diego, California
Citation
HAN, Xianjing; ZHU, Bin; HU, Shiqi; LI, Franklin Mingzhe; CARRINGTON, Patrick; ZIMMERMANN, Roger; and CHEN, Jingjing.
OSCBench: Benchmarking object state change in text-to-video generation. (2026). Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL 2026), San Diego, California, July 2-7. 30867-30884.
Available at: https://ink.library.smu.edu.sg/sis_research/11173
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Additional URL
https://doi.org/10.18653/v1/2026.acl-long.1425
Included in
Artificial Intelligence and Robotics Commons, Graphics and Human Computer Interfaces Commons