Geo-VLA:内化地图语义的几何感知 VLA 规划
Original title: Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics
IndustryAI 67
Source: arXiv cs.ROPublish time unverified
arXiv:2608.21440v1 Announce Type: new Abstract: Vision-language-action (VLA) models have advanced end-to-end autonomous driving by leveraging foundation models for semantic reasoning and long-tail generalization. However, their planning performance remains limited in complex driving environments because image-only representations inadequately capture planning-relevant road geometry and topology.