Publication Type
Conference Proceeding Article
Version
publishedVersion
Publication Date
6-2026
Abstract
Vision-language segmentation models such as SAM3 enable flexible, prompt-driven visual grounding, but inherit large, general-purpose text encoders originally designed for open-ended language understanding. In practice, segmentation prompts are short, structured, and semantically constrained, leading to substantial over-provisioning in text encoder capacity and persistent computational and memory overhead. In this paper, we perform a large-scale anatomical analysis of text prompting in vision–language segmentation, covering 404,796 real prompts across multiple benchmarks. Our analysis reveals severe redundancy: most context windows are underutilized, vocabulary usage is highly sparse, and text embeddings lie on a low-dimensional manifold despite high-dimensional representations. Motivated by these findings, we propose SAM3-LiteText, a lightweight text encoding framework that replaces the original SAM3 text encoder with a compact MobileCLIP student that is optimized by knowledge distillation. Extensive experiments on image and video segmentation benchmarks show that SAM3-LiteText reduces text encoder parameters by up to 88%, substantially reducing static memory footprint, while maintaining segmentation performance comparable to the original model. Code: https://github.com/SimonZeng7108/efficientsam3/tree/sam3_litetext.
Keywords
Vision-Language Models, Image Segmentation, Model Compression, Multimedia content extraction, Knowledge Distillation
Discipline
Artificial Intelligence and Robotics | Graphics and Human Computer Interfaces
Research Areas
Intelligent Systems and Optimization
Areas of Excellence
Digital transformation
Publication
ICMR '26: Proceedings of the 2026 International Conference on Multimedia Retrieval, Amsterdam, The Netherlands, June 16-19
First Page
1147
Last Page
1156
ISBN
9798400726170
Identifier
10.1145/3805622.3810586
Publisher
ACM
City or Country
New York
Citation
ZENG, Chengxi; JIANG, Yuxuan; GAO, Ge; WANG, Shuai; DANIER, Duolikun; ZHU, Bin; RUDINAC, Stevan; BULL, David; and ZHANG, Fan.
SAM3-LiteText: An anatomical study of the SAM3 text encoder for efficient vision-language segmentation. (2026). ICMR '26: Proceedings of the 2026 International Conference on Multimedia Retrieval, Amsterdam, The Netherlands, June 16-19. 1147-1156.
Available at: https://ink.library.smu.edu.sg/sis_research/11129
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Additional URL
https://doi.org/10.1145/3805622.3810586
Included in
Artificial Intelligence and Robotics Commons, Graphics and Human Computer Interfaces Commons