Publication Type
Conference Proceeding Article
Version
publishedVersion
Publication Date
1-2026
Abstract
Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-label methods rely on a fixed set of textual proxies, which (i) sparsely sample the semantic space beyond in-distribution (ID) classes and (ii) remain static while only visual features drift, leading to cross-modal misalignment and unstable predictions. In this paper, we propose CoEvo, a training- and annotation-free test-time framework that performs bidirectional, sample-conditioned adaptation of both textual and visual proxies. Specifically, CoEvo introduces a proxy-aligned co-evolution mechanism to maintain two evolving proxy caches, which dynamically mines contextual textual negatives guided by test images and iteratively refines visual proxies, progressively realigning cross-modal similarities and enlarging local OOD margins. Finally, we dynamically re-weight the contributions of dual-modal proxies to obtain a calibrated OOD score that is robust to distribution shift. Extensive experiments on standard benchmarks demonstrate that CoEvo achieves state-of-the-art performance, improving AUROC by 1.33% and reducing FPR95 by 45.98% on ImageNet-1K compared to strong negative-label baselines.
Discipline
Artificial Intelligence and Robotics | Graphics and Human Computer Interfaces
Research Areas
Intelligent Systems and Optimization
Publication
Proceedings of the AAAI Conference on Artificial Intelligence: Singapore, January 20-27
Volume
40
First Page
15770
Last Page
15778
ISBN
9781577359067
Identifier
10.1609/aaai.v40i18.38608
Publisher
AAAI
City or Country
Singapore
Citation
Tang, Hao; LIU, Yu; YAN, Shuanglin; Shen, Fei; HE, Shengfeng; and Qin, Jing.
Cross-modal proxy evolving for OOD detection with Vision-Language Models. (2026). Proceedings of the AAAI Conference on Artificial Intelligence: Singapore, January 20-27. 40, 15770-15778.
Available at: https://ink.library.smu.edu.sg/sis_research/11200
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Included in
Artificial Intelligence and Robotics Commons, Graphics and Human Computer Interfaces Commons