Publication Type

Journal Article

Version

acceptedVersion

Publication Date

1-2026

Abstract

Multimodal sentiment models often become over-reliant on the “easiest” modality (typically text), leading to three coupled sub-problems: (i) representation-level dominance, where weaker modalities contribute little to the fused representation; (ii) optimization-level dominance, where the strongest modality drives most gradient updates and suppresses learning in others; and (iii) robustness degradation, where audio or vision fail under noise or missing inputs at test time. We present TEMPO, a plug-and-play training framework that mitigates these issues by rebalancing learning pressure across modalities while leaving inference unchanged. For each mini-batch, TEMPO estimates relative modality strength and applies two synchronized, training-only controls: selective forward attenuation and backward gradient equilibration. On IEMOCAP and MELD, TEMPO improves accuracy and weighted F1 over strong multimodal baselines, increases the standalone usefulness of weaker modalities, and offers higher robustness under missing- or corrupted-modality stress. Across three benchmarks, TEMPO improves accuracy by 3.2–6.7% and reduces calibration error by 18–34%, demonstrating consistent gains in both performance and reliability with negligible computational overhead.

Keywords

Multimodal sentiment analysis, Modality imbalance, Training-time modulation, Adaptive attenuation

Discipline

Artificial Intelligence and Robotics

Publication

IEEE Transactions on Affective Computing

First Page

1

Last Page

17

ISSN

1949-3045

Identifier

10.1109/TAFFC.2026.3657064

Publisher

Institute of Electrical and Electronics Engineers

Additional URL

https://doi.org/10.1109/TAFFC.2026.3657064

Share

COinS