Publication Type
Conference Proceeding Article
Version
publishedVersion
Publication Date
6-2026
Abstract
Vision-and-Language Navigation in Continuous Environment (VLN-CE) requires an agent to follow language instructions to navigate the target destination. With the advancement of large language models (LLMs), recent efforts have explored adapting them for zero-shot VLN-CE, offering a promising solution in addressing the drawbacks of poor generalization in the training-based paradigm. However, existing LLM-based works primarily perform naive reasoning for decision-making and lack feedback, e.g., reviewing historical errors and predicting future potentials. Consequently, it may suffer from continuous failure for those initial error tasks. In this paper, we rethink LLM-based zero-shot VLN-CE and propose a new paradigm, named EvoNav, to improve the agent's decision-making with future thought and history experience via Future Chain-of-Thought (F-CoT) and History Chain-of-Experience (H-CoE). F-CoT predicts future actions and landmarks as thoughts to assist navigation progress estimation and direction selection, while H-CoE summarizes historical trajectories and scenes as experience to improve navigation decision reliability. Both F-CoT and H-CoE cooperatively evolve the agent’s decision-making. Extensive experiments in both the simulator and real-world environments demonstrate the effectiveness of our EvoNav. Source code will be released.
Discipline
Artificial Intelligence and Robotics | Graphics and Human Computer Interfaces
Research Areas
Software and Cyber-Physical Systems
Areas of Excellence
Digital transformation
Publication
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026), Denver, Colorado, United States, June 3-7
First Page
15177
Last Page
15187
Publisher
IEEE Computer Society
City or Country
Los Alamitos, CA
Citation
DAI, Guangzhao; WANG, Shuo; WANG, Zihan; XIE, Guo-Sen; YANG, Yang; PAN, Jinshan; SUN, Qianru; and SHU, Xiangbo.
History to future: Evolving agent with experience and thought for zero-shot vision-and-language navigation. (2026). Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2026), Denver, Colorado, United States, June 3-7. 15177-15187.
Available at: https://ink.library.smu.edu.sg/sis_research/11199
Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-No Derivative Works 4.0 International License.
Additional URL
https://cvpr.thecvf.com/virtual/2026/poster/39872
Included in
Artificial Intelligence and Robotics Commons, Graphics and Human Computer Interfaces Commons