Publication Type
Journal Article
Version
submittedVersion
Publication Date
4-2026
Abstract
Recent approaches attempt to adapt powerful interactive segmentation models, such as SAM, to interactive matting and fine-tune the models based on synthetic matting datasets. However, models trained on synthetic data fail to generalize to complex and occlusion scenes. We address this challenge by proposing a new matting dataset based on the COCO dataset, namely COCO-Matting. It selects real-world complex images from COCO and converts semantic segmentation masks to matting labels. The built COCO-Matting comprises an extensive collection of 36,980 human instance-level alpha mattes in complex natural scenarios. Furthermore, existing SAM-based matting methods extract intermediate features and masks from a frozen SAM and only train a lightweight matting decoder by end-to-end matting losses, which do not fully exploit the potential of the pre-trained SAM. Thus, we propose SEMat which revamps the network architecture and training objectives. For network architecture, the proposed feature-aligned transformer learns to extract fine-grained edge and transparency features. The proposed matte-aligned decoder aims to segment matting-specific objects and convert coarse masks into high-precision mattes. For training objectives, the proposed regularization and trimap loss aim to retain the prior from the pre-trained model and push the matting logits extracted from the mask decoder to contain trimap-based semantic information. Extensive experiments across seven diverse datasets demonstrate the superior performance of our method, proving its efficacy in interactive natural image matting. Code is available at https://github.com/XiaRho/SEMat
Discipline
Artificial Intelligence and Robotics | Graphics and Human Computer Interfaces
Research Areas
Intelligent Systems and Optimization
Areas of Excellence
Digital transformation
Publication
IEEE Transactions on Circuits and Systems for Video Technology
Volume
36
Issue
4
ISSN
1051-8215
Identifier
10.1109/TCSVT.2025.3637212
Publisher
Institute of Electrical and Electronics Engineers
Citation
XIA, Ruihao; LIANG, Yu; JIANG, Peng-Tao; ZHANG, Hao; SUN, Qianru; TANG, Yang; LI, Bo; and Pan ZHOU.
SEMat: Semantic enhanced natural image interactive matting. (2026). IEEE Transactions on Circuits and Systems for Video Technology. 36, (4),.
Available at: https://ink.library.smu.edu.sg/sis_research/11180
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Additional URL
https://doi.org/10.1109/TCSVT.2025.3637212
Included in
Artificial Intelligence and Robotics Commons, Graphics and Human Computer Interfaces Commons