Publication Type

Conference Proceeding Article

Version

publishedVersion

Publication Date

6-2026

Abstract

Frozen Large Video Language Models (LVLMs) are increasingly employed in micro-video recommendation (MVR) due to their strong multimodal understanding. However, existing apporches typically deploy LVLMs as fixed black-box feature extractors without systematically comparing alternative representation strategies. To address this gap, we present the first systematic empirical study on various feature extraction paradigms and integration strategies, along with hierarchical representations from frozen LVLMs for MVR. Extensive experiments on representative LVLMs reveal that hidden states from multiple decoder layers provide richer and more effective representations for MVR. Guided by this insight, we propose the Dual Feature Fusion (DFF) Framework, a lightweight approach that adaptively fuses multi-layer representations from frozen LVLMs with ID embeddings. DFF achieves state-of-the-art performance on two real-world micro-video recommendation benchmarks, consistently outperforming strong baselines and providing a principled approach to integrating off-the-shelf large vision-language models into micro-video recommender systems.

Keywords

Micro-video Recommendation, Large Video Language Model, Feature Fusion

Discipline

Artificial Intelligence and Robotics

Research Areas

Intelligent Systems and Optimization

Areas of Excellence

Digital transformation

Publication

ICMR '26: Proceedings of the 2026 International Conference on Multimedia Retrieval, Amsterdam, The Netherlands, June 16-19

First Page

40

Last Page

49

ISBN

9798400726170

Identifier

10.1145/3805622.3810736

Publisher

ACM

City or Country

New York

Additional URL

https://doi.org/10.1145/3805622.3810736

Share

COinS