STPrompt++: Prompting vision-language models for weakly supervised video anomaly detection and fine-grained localization

Publication Type

Journal Article

Version

publishedVersion

Publication Date

7-2028

Abstract

Traditional weakly supervised video anomaly detection (WSVAD) tasks typically rely on coarse-grained frame-level labels for training. Although this approach reduces annotation costs, it results in weak semantic understanding and spatial localization capabilities due to the absence of fine-grained annotations, hindering precise pixel-level anomaly detection and localization. Thanks to the success of vision-language models (VLMs), e.g., CLIP, recent approaches leveraging large VLMs focus on exploiting their strong semantic understanding capabilities, but they typically feed only keyframes or short video segments into the models, without supplying sufficient prior contextual information (e.g., contextual frames around anomalies, zoomed-in anomaly regions, and detailed anomaly descriptions), which restricts the models’ capability for fine-grained anomaly understanding and precise localization. More recently, a few methods leveraging VLMs, attempt to achieve training-free spatial anomaly localization by fusing patch-level visual features with simple textual features. However, these methods employ simplistic textual descriptions, lacking deep semantic comprehension of anomalies, leading to coarse localization results with significant irrelevant background noise. To address these issues, we propose STPrompt, a novel weakly supervised spatio-temporal video anomaly detection and localization method based on VLMs. In our work, we systematically leverage preliminary coarse localization regions derived from anomaly scores as spatial priors, together with contextual frames around keyframes, zoomed-in views of suspected anomalous regions, and refined textual descriptions of anomalies. This comprehensive prompting mechanism guides the VLMs toward deep semantic comprehension of video anomalies, enabling accurate pixel-level spatial localization. The proposed STPrompt requires no additional training and significantly enhances the precision of anomaly understanding and localization through a carefully designed multi-round and multi-modal prompting mechanism. Extensive experiments on two widely used WSVAD benchmarks, UCF-Crime and UBnormal, show that our method achieves state-of-the-art spatial localization performance and competitive temporal anomaly detection results. Notably, on UCF-Crime dataset, our approach improves spatial localization accuracy (in TIoU) by 5.61% over the current best method (from 23.90% to 29.51%), underscoring its superior capabilities in precise anomaly localization and semantic understanding.

Keywords

video anomaly detection, anomaly localization, vision-language models

Discipline

Artificial Intelligence and Robotics | Graphics and Human Computer Interfaces

Research Areas

Intelligent Systems and Optimization

Areas of Excellence

Digital transformation

Publication

ACM Transactions on Multimedia Computing, Communications and Applications

Volume

22

Issue

8

First Page

1

Last Page

22

ISSN

1551-6857

Identifier

10.1145/3828722

Publisher

Association for Computing Machinery (ACM)

Additional URL

https://doi.org/10.1145/3828722

This document is currently not available here.

Share

COinS