Publication Type
Conference Proceeding Article
Version
publishedVersion
Publication Date
7-2026
Abstract
Personalized generative recommender systems have emerged as a promising solution for fashion recommendation. However, existing methods primarily rely on implicit visual embeddings from historical interactions, which often contain preference-irrelevant information and result in insufficient user behavior modeling. Moreover, these models typically generate only item images, providing limited interpretability. To address these limitations, we propose DualFashion, a Dual-Diffusional Generative Fashion Recommendation Architecture that jointly models image and text modalities for personalized and explainable recommendation. DualFashion adopts a dual-diffusion Transformer with image and text branches, where structured attribute-level captions and visual outfit information are jointly used as conditioning signals to model user behavior. The proposed architecture produces both fashion item images and textual descriptions, ensuring visual compatibility while providing explicit semantic interpretability. Furthermore, we introduce a text-augmented fine-tuning strategy that enhances generation diversity and enables effective cross-modal knowledge transfer without incurring heavy computational costs. Extensive experiments on iFashion and Polyvore-U across Personalized Fill-in-the-Blank and Generative Outfit Recommendation tasks demonstrate that DualFashion achieves strong performance in behavior modeling, interpretability, and efficiency compared to state-of-the-art methods. Our code and model checkpoints are available at https://github.com/LinkMingzhe/DualFashion.
Keywords
Fashion Outfit Generation, Fashion Image Generation, GenerativeFashion Recommendation
Discipline
Artificial Intelligence and Robotics | Graphics and Human Computer Interfaces
Research Areas
Intelligent Systems and Optimization
Areas of Excellence
Digital transformation
Publication
SIGIR '26: Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, Melbourne, Australia, July 20-24
First Page
2264
Last Page
2273
Identifier
10.1145/3805712.3809557
Publisher
ACM
City or Country
New York
Citation
YU, Mingzhe; WU, Lei; SUN, Qianru; and MA, Yunshan.
Dual-diffusional generative fashion recommendation. (2026). SIGIR '26: Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, Melbourne, Australia, July 20-24. 2264-2273.
Available at: https://ink.library.smu.edu.sg/sis_research/11273
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Additional URL
https://doi.org/10.1145/3805712.3809557
Included in
Artificial Intelligence and Robotics Commons, Graphics and Human Computer Interfaces Commons