Publication Type
Journal Article
Version
acceptedVersion
Publication Date
9-2026
Abstract
Foundation models like CLIP allow zero-shot transfer on various tasks without additional training data. Yet, the zero-shot performance is less competitive than a fully supervised one. Thus, fine-tuning and ensembling are also commonly adopted to better fit the downstream tasks. However, we argue that such prior work has overlooked the inherent biases in foundation models. Due to the highly imbalanced Web-scale training set, foundation models are inevitably skewed toward frequent semantics, and thus the subsequent fine-tuning or ensembling is still biased. In this study, we systematically examine the biases in foundation models and demonstrate the efficacy of our proposed Generalized Logit Adjustment (GLA) method. Note that bias estimation in foundation models is challenging, as most pre-train data cannot be explicitly accessed like in traditional long-tailed classification tasks. To this end, GLA offers two alternative methods for debiasing: the first is an optimization-based bias estimation built on Bayes optimal criterion, and the second identifies label bias through an eigenvector derived from a matrix of zero-shot predictions. As our work resolves a fundamental flaw in the pre-training, the proposed GLA demonstrates significant improvements across a diverse range of tasks: it achieves 1.5 pp accuracy gains on ImageNet, a large average improvement (1.9−4.4 pp) on 11 few-shot datasets, 2.4 pp gains on long-tailed classification.
Keywords
Vision-language models, Label distribution bias, Image classification, Long-tail learning
Discipline
Artificial Intelligence and Robotics | Graphics and Human Computer Interfaces
Research Areas
Intelligent Systems and Optimization
Areas of Excellence
Digital transformation
Publication
International Journal of Computer Vision
Volume
134
First Page
1
Last Page
20
ISSN
0920-5691
Identifier
10.1007/s11263-026-03012-w
Publisher
Springer
Citation
ZHU, Beier; SUN, Qianru; YANG, Xun; and ZHANG, Hanwang.
Generalized logit adjustment: Improved fine-tuning by mitigating label bias in zero-shot vision models. (2026). International Journal of Computer Vision. 134, 1-20.
Available at: https://ink.library.smu.edu.sg/sis_research/11290
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Additional URL
https://doi.org/10.1007/s11263-026-03012-w
Included in
Artificial Intelligence and Robotics Commons, Graphics and Human Computer Interfaces Commons