Publication Type
Conference Proceeding Article
Version
publishedVersion
Publication Date
6-2026
Abstract
Large Multi-modal Models (LMMs) have significantly advanced a variety of vision-language tasks. The scalability and availability of high-quality training data play a pivotal role in the success of LMMs. In the realm of food, while comprehensive food datasets such as Recipe1M offer an abundance of ingredient and recipe information, they often fall short of providing ample data for nutritional analysis. The Recipe1M+ dataset, despite offering a subset for nutritional evaluation, is limited in the scale and accuracy of nutrition information. To bridge this gap, we introduce Uni-Food, a unified food dataset that comprises over 100,000 images with various food labels, including categories, ingredients, recipes, and ingredient-level nutritional information. To mitigate the conflicts arising from multi-task supervision during fine-tuning of LMMs, we introduce a novel Linear Rectification Mixture of Diverse Experts (RoDE) approach. RoDE utilizes a diverse array of experts to address tasks of varying complexity, thereby facilitating the coordination of trainable parameters, i.e., it allocates more parameters for more complex tasks and, conversely, fewer parameters for simpler tasks. RoDE implements linear rectification union to refine the router’s functionality, thereby enhancing the efficiency of sparse task allocation. These design choices endow RoDE with features that ensure GPU memory efficiency and ease of optimization. Extensive experiments validate the effectiveness of our approach in addressing the inherent challenges of food-related multitasking. UniFood Project
Keywords
Food Computing, Large Multi-Modal Models, Mixture of Experts, Heterogeneous Experts
Discipline
Artificial Intelligence and Robotics | Databases and Information Systems
Research Areas
Intelligent Systems and Optimization
Areas of Excellence
Digital transformation
Publication
ICMR '26: Proceedings of the 2026 International Conference on Multimedia Retrieval, Amsterdam, The Netherlands, June 16-19
First Page
2457
Last Page
2466
ISBN
9798400726170
Identifier
10.1145/3805622.3810616
Publisher
ACM
City or Country
New York
Citation
JIAO, Pengkun; WU, Xinlan; ZHU, Bin; CHEN, Jingjing; NGO, Chong-wah; and YU-GANG.
RoDE: Linear rectified mixture of diverse experts for food large multi-modal models. (2026). ICMR '26: Proceedings of the 2026 International Conference on Multimedia Retrieval, Amsterdam, The Netherlands, June 16-19. 2457-2466.
Available at: https://ink.library.smu.edu.sg/sis_research/11126
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Additional URL
https://doi.org/10.1145/3805622.3810616