Publication Type
Conference Proceeding Article
Version
acceptedVersion
Publication Date
9-2026
Abstract
Vision-Language-Action (VLA) models enable robots to perform manipulation tasks directly from natural language instructions and are increasingly viewed as a foundation for generalist robotic policies. However, their reliability under Out-Of-Distribution (OOD) instructions remains underexplored. In this paper, we reveal a critical failure mode in which VLA policies continue executing visually plausible actions even when the language instruction contradicts the scene. We refer to this phenomenon as linguistic blindness, where VLA policies prioritize visual priors over instruction semantics during action generation. To systematically analyze this issue, we introduce ICBench, a diagnostic benchmark constructed from the LIBERO dataset that probes language–action coupling by injecting controlled OOD instruction contradictions while keeping the visual environment unchanged. Evaluations on three representative VLA architectures, including π0, π0.5, and OpenVLA-OFT, show that these models frequently succeed at tasks despite logically impossible instructions, revealing a strong visual bias in action generation. To mitigate this issue, we propose Instruction-Guided Attention Recalibration (IGAR), a train-free inference-time mechanism that rebalances attention distributions to restore the influence of language instructions. IGAR operates without retraining or architectural modification and can be directly applied to existing VLA models. Experiments across 30 LIBERO tasks demonstrate that IGAR substantially reduces erroneous execution under OOD contradictory instructions while preserving baseline task performance. We additionally validate the approach on a real Franka robotic arm, where IGAR effectively prevents manipulation triggered by inconsistent instructions.
Keywords
Vision-Language-Action Models, Robotic Manipulation, Attention Recalibration
Discipline
Artificial Intelligence and Robotics | Graphics and Human Computer Interfaces
Research Areas
Intelligent Systems and Optimization
Areas of Excellence
Digital transformation
Publication
Proceedings of the 19th European Conference on Computer Vision (ECCV 2026), Malmö, Sweden, September 8-12
Last Page
1
Publisher
Springer
City or Country
Cham
Citation
ZHANG, Ninghao; ZHU, Bin; ZHOU, Shijie; and CHEN, Jingjing.
Restoring linguistic grounding in VLA models via train-free attention recalibration. (2026). Proceedings of the 19th European Conference on Computer Vision (ECCV 2026), Malmö, Sweden, September 8-12. 1.
Available at: https://ink.library.smu.edu.sg/sis_research/11175
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Included in
Artificial Intelligence and Robotics Commons, Graphics and Human Computer Interfaces Commons