1 Tsinghua University2 Pengcheng Laboratory3 South China University of Technology4 International Digital Economy Academy5 Ping An Technology (Shenzhen) Co., Ltd. Shenzhen, China6 Harbin Institute of Technology (Shenzhen)
Progress Myopia. We identify a failure mode in vision-language navigation: agents may continue acting confidently when the observations needed to ground task progress are missing. An analysis of successful and failed segments shows that decision confidence does not reliably reflect this uncertainty.
SeekVLN. Our framework couples semantic progress reasoning with active acquisition of task-relevant views. FRG builds evidence-seeking supervision from offline expert trajectories; C2PO compares seeking with direct navigation from the same state to learn when the additional evidence improves progress.
Empirical validation. On R2R-CE and RxR-CE validation-unseen, the full model improves success rate over the base model by 12.7 and 7.5 percentage points, respectively. Simulated and real-world rollouts illustrate targeted evidence seeking.
Real-world rollout
At the end of a hallway, the destination chair is absent from the forward view and the instruction does not specify a turn direction. SeekVLN acquires a side view, locates the chair to its left, and completes the task.
Real-world rolloutUnitree Go2
Robot view
Acquired evidence / left view
Chair located
Evidence seeking in the real world.At the hallway’s end, the chair is outside the forward view. Looking left reveals the evidence needed for the next turn.
Vision-language navigation requires an agent to track completed subgoals and determine where to go next over a long instruction. Partial egocentric observations and ambiguous instructions can leave the evidence required for that decision unavailable: a landmark may lie outside the field of view, or a junction may admit multiple plausible turns.
Existing agents largely reason from the observations already available. In the paper, three representative models show similar decision confidence on successful and failed trajectory segments, even when navigation has deviated from the goal. We call this failure to recognize unreliable progress grounding Progress Myopia. SeekVLN addresses it by acquiring task-relevant observations before committing to a navigation action.
Figure 1. Illustration of Progress Myopia. Missing landmarks and ambiguous instructions make the current observation insufficient; unreliable navigation decisions can nevertheless remain highly confident.
Methodology
At each decision, SeekVLN selects NAV to act from the available context or SEEK to acquire supplementary left, front, and right views. After seeking, it identifies key evidence, reasons about completed and upcoming subgoals, and predicts the navigation action. Two training stages learn this behavior.
Stage 1 · Supervised fine-tuning
FRG Future-guided Reverse Generation
FRG uses future expert actions to select where seeking may help. A VLM then annotates the key visual evidence and task progress, turning offline expert trajectories into supervision for evidence seeking and progress reasoning without additional expert interaction.
C2PO compares a SEEK branch with a direct-NAV branch from the same state. A contrastive reward measures the difference in subsequent progress and, together with an outcome reward, assigns credit to useful seeking decisions.
Figure 2. Overview of SeekVLN. The policy decides whether to navigate or seek evidence. FRG constructs supervision from offline trajectories, and C2PO jointly optimizes seeking and navigation.
Counterfactual branches are sampled during training; evaluation uses a single policy.
Results
On R2R-CE and RxR-CE validation-unseen, FRG-SFT raises success rate from 54.8 to 61.0 and from 52.2 to 55.7, respectively. C2PO-RFT further raises it to 67.5 and 59.7. The full two-stage model gains 12.7 and 7.5 percentage points over the base model.
On a separate 100-episode R2R-CE subset, adaptive seeking reaches 73% success at a 29.8% seek rate; periodic seeking reaches 62% success while seeking at 50% of decisions. During reinforcement fine-tuning, beneficial action changes after seeking increase from 45.6% to 55.9%.
Figure 3. Deep analysis of evidence-seeking behavior, reproduced from the paper.
Qualitative navigation
In a simulated trajectory, SeekVLN first locates a dining area to the left and later finds a hallway beyond the kitchen island to the right. Each supplementary view resolves uncertainty about instruction progress before the next action.
Figure 4. Two SEEK decisions in a simulated trajectory. The agent first locates the dining area on the left, then identifies a hallway beyond the kitchen island on the right.
@misc{seekvln2026,
title = {Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation},
author = {Wang, Zhimin and Zhu, Meiyuan and Wu, Duo and Kang, Linjia and Wang, Yajun and Ni, Yuan and Wang, Xiaohang and Pan, Tianlu and Jiang, Jingyan and Wang, Yaowei and Wang, Zhi},
year = {2026},
eprint = {2609.37353},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.37353}
}