Embodied Intelligence Observer

DeicticVLA:统一语言与指示手势指令模式的 VLA

Original title: DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA

ResearchAI 75

Source: arXiv cs.ROPublish time unverified

arXiv:2608.28108v1 Announce Type: new Abstract: Vision-Language-Action models (VLAs) allow users to specify manipulation tasks in natural language, but distinguishing a target or placement goal among objects of the same category or similar appearance requires detailed expressions that VLAs may not use reliably.