WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN
Research
Source: arXiv cs.ROPublish time unverified
arXiv:2608.07267v2 Announce Type: replace-cross Abstract: Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly to navigation actions. Although semantically capable, such action-centric training does not explicitly model how the agent's visual observations should evolve under its predicted motion.