具身智能观察

Mind-VLA:指令感知的空间表征对齐

原标题:Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models

技术动态AI 80

来源:arXiv cs.RO发布时间待核实

arXiv:2608.04633v2 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) methods improve generalization by aligning their representations with 3D scene geometry. However, these methods are fundamentally instruction-agnostic: the representations align the entire scene uniformly, neglecting the 3D geometry of the specific target object designated by the language instruction.

Mind-VLA:指令感知的空间表征对齐 | 具身智能观察