SeekVLN / Research paper2026
SeekVLN

Seek Before You Move:Evidence Seeking for Progress Grounding
in Vision-Language Navigation

Zhimin Wang1,2 Meiyuan Zhu1 Duo Wu1,2 Linjia Kang1 Yajun Wang3,4 Yuan Ni5 Xiaohang Wang5 Tianlu Pan2 Jingyan Jiang1 Yaowei Wang6,2,* Zhi Wang1,*
1 Tsinghua University2 Pengcheng Laboratory3 South China University of Technology4 International Digital Economy Academy5 Ping An Technology (Shenzhen) Co., Ltd. Shenzhen, China6 Harbin Institute of Technology (Shenzhen)
* Corresponding authors: Yaowei Wang and Zhi Wang

Contributions

Progress Myopia. We identify a failure mode in vision-language navigation: agents may continue acting confidently when the observations needed to ground task progress are missing. An analysis of successful and failed segments shows that decision confidence does not reliably reflect this uncertainty.

SeekVLN. Our framework couples semantic progress reasoning with active acquisition of task-relevant views. FRG builds evidence-seeking supervision from offline expert trajectories; C2PO compares seeking with direct navigation from the same state to learn when the additional evidence improves progress.

Empirical validation. On R2R-CE and RxR-CE validation-unseen, the full model improves success rate over the base model by 12.7 and 7.5 percentage points, respectively. Simulated and real-world rollouts illustrate targeted evidence seeking.

Real-world rollout

At the end of a hallway, the destination chair is absent from the forward view and the instruction does not specify a turn direction. SeekVLN acquires a side view, locates the chair to its left, and completes the task.

Real-world rolloutUnitree Go2
A Unitree Go2 robot turns at the end of a hallway to seek evidence.Robot view
The robot's left view reveals the destination chair.Acquired evidence / left view
Chair located
Evidence seeking in the real world.At the hallway’s end, the chair is outside the forward view. Looking left reveals the evidence needed for the next turn.

Qualitative example from the paper. View the complete rollout figure ↗

Introduction

Vision-language navigation requires an agent to track completed subgoals and determine where to go next over a long instruction. Partial egocentric observations and ambiguous instructions can leave the evidence required for that decision unavailable: a landmark may lie outside the field of view, or a junction may admit multiple plausible turns.

Existing agents largely reason from the observations already available. In the paper, three representative models show similar decision confidence on successful and failed trajectory segments, even when navigation has deviated from the goal. We call this failure to recognize unreliable progress grounding Progress Myopia. SeekVLN addresses it by acquiring task-relevant observations before committing to a navigation action.

Figure 01 / Progress MyopiaOpen ↗
Figure 1 from the paper: out-of-view landmarks, ambiguous instructions, and decision-confidence analysis of Progress Myopia.
Figure 1. Illustration of Progress Myopia. Missing landmarks and ambiguous instructions make the current observation insufficient; unreliable navigation decisions can nevertheless remain highly confident.

Methodology

At each decision, SeekVLN selects NAV to act from the available context or SEEK to acquire supplementary left, front, and right views. After seeking, it identifies key evidence, reasons about completed and upcoming subgoals, and predicts the navigation action. Two training stages learn this behavior.

Stage 1 · Supervised fine-tuning

FRG Future-guided Reverse Generation

FRG uses future expert actions to select where seeking may help. A VLM then annotates the key visual evidence and task progress, turning offline expert trajectories into supervision for evidence seeking and progress reasoning without additional expert interaction.

Stage 2 · Reinforcement fine-tuning

C2PO Counterfactual Contrastive Policy Optimization

C2PO compares a SEEK branch with a direct-NAV branch from the same state. A contrastive reward measures the difference in subsequent progress and, together with an outcome reward, assigns credit to useful seeking decisions.

Figure 02 / SeekVLN frameworkOpen ↗
Figure 2: SeekVLN dual-mode navigation, FRG supervision, and C2PO counterfactual branch comparison.
Figure 2. Overview of SeekVLN. The policy decides whether to navigate or seek evidence. FRG constructs supervision from offline trajectories, and C2PO jointly optimizes seeking and navigation.

Counterfactual branches are sampled during training; evaluation uses a single policy.

Results

On R2R-CE and RxR-CE validation-unseen, FRG-SFT raises success rate from 54.8 to 61.0 and from 52.2 to 55.7, respectively. C2PO-RFT further raises it to 67.5 and 59.7. The full two-stage model gains 12.7 and 7.5 percentage points over the base model.

Table 01 / Benchmark resultsOpen ↗
Original Table 1 comparing 20 methods and all reported R2R-CE and RxR-CE validation-unseen metrics. SeekVLN-C2PO-RFT achieves R2R NE 3.7, OSR 75.2, SR 67.5, SPL 61.4; RxR NE 4.9, SR 59.7, SPL 50.3, nDTW 63.6.
Original Table 1 from the manuscript. View in the paper ↗

Adaptive evidence seeking

On a separate 100-episode R2R-CE subset, adaptive seeking reaches 73% success at a 29.8% seek rate; periodic seeking reaches 62% success while seeking at 50% of decisions. During reinforcement fine-tuning, beneficial action changes after seeking increase from 45.6% to 55.9%.

Figure 03 / Evidence-seeking behaviorOpen ↗
Original paper Figure 3: trigger strategy comparison, beneficial action change rate during training, and progress gain from evidence seeking.
Figure 3. Deep analysis of evidence-seeking behavior, reproduced from the paper.

Qualitative navigation

In a simulated trajectory, SeekVLN first locates a dining area to the left and later finds a hallway beyond the kitchen island to the right. Each supplementary view resolves uncertainty about instruction progress before the next action.

Figure 04 / Qualitative rolloutsOpen ↗
Original simulation rollout figure showing evidence seeking for a dining area and a hallway beyond a kitchen island.
Figure 4. Two SEEK decisions in a simulated trajectory. The agent first locates the dining area on the left, then identifies a hallway beyond the kitchen island on the right.

Cite the arXiv preprint.

seekvln.bib
@misc{seekvln2026,
  title = {Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation},
  author = {Wang, Zhimin and Zhu, Meiyuan and Wu, Duo and Kang, Linjia and Wang, Yajun and Ni, Yuan and Wang, Xiaohang and Pan, Tianlu and Jiang, Jingyan and Wang, Yaowei and Wang, Zhi},
  year  = {2026},
  eprint = {2609.37353},
  archivePrefix = {arXiv},
  url   = {https://arxiv.org/abs/2609.37353}
}