具身智能观察

动作与语言条件化的具身控制视频评估

原标题:Action- and Language-Conditioned Video Assessment for Embodied Control

技术动态AI 70

来源:arXiv cs.RO发布时间待核实

arXiv:2608.08273v2 Announce Type: replace Abstract: Vision-based embodied agents executing multi-step natural language instructions require feedback mechanisms that assess task progress over complete trajectories. Conventional approaches based on final-frame matching or continuous embedding similarity may overlook intermediate transitions that are necessary for determining whether an instruction has been completed.