具身智能观察

GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model

技术动态

来源:arXiv cs.RO发布时间待核实

arXiv:2608.24959v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models encode visual observations as flat 2D patch tokens that carry no intrinsic geometric structure, and augmenting them with dense monocular depth injects per-pixel scalar values that encode neither surface orientation nor geometric confidence. This leaves the policy with limited structured spatial reasoning for action prediction.