Publication Type

Conference Proceeding Article

Version

publishedVersion

Publication Date

6-2026

Abstract

We propose InterFold, a framework for learning and applying interpretable semantic manifolds in latent diffusion models, without requiring binary or paired supervision. Existing methods for semantic editing either rely on limited paired data or uncover only coarse, unsupervised directions that fail to capture user-specific, fine-grained attributes. InterFold addresses these limitations by learning a target attribute manifold in the H-space of diffusion models using only a set of positive, unlabeled examples. To edit a new image, InterFold projects its H-space representation toward this learned manifold through test-time optimization, enabling precise, identity-preserving modifications of complex, non-binary concepts. To make these edits effective in modern latent diffusion models, we introduce the Manifold Adapter, a lightweight cross-attention module that transfers semantic intent from edited H-space codes into the generative latent space, without altering the pretrained model. Extensive experiments demonstrate that InterFold achieves superior edit accuracy and identity consistency compared to existing methods, offering a flexible and interpretable solution for high-fidelity semantic image editing.

Keywords

Interpretable Semantic Editing, Latent Diffusion Models, Positive-Only Representation Learning

Discipline

Artificial Intelligence and Robotics

Research Areas

Intelligent Systems and Optimization

Areas of Excellence

Digital transformation

Publication

ICMR '26: Proceedings of the 2026 International Conference on Multimedia Retrieval, Amsterdam, The Netherlands, June 16-19

First Page

1861

Last Page

1869

ISBN

9798400726170

Publisher

ACM

City or Country

New York

Additional URL

https://doi.org/10.1145/3805622.3810585

Share

COinS