Embodied Intelligence Observer

跨视角一致性只约束动作块:世界动作模型对未见机位成功率提升 12 个百分点

Original title: Selective Cross-View Consistency for World Action Models: Held-Out Viewpoint Robustness Without Test-Time Camera Information

CovariantIndustryAI 85

Source: arXiv cs.ROPublish time unverified

针对世界动作模型(WAM)的相机视角扰动,研究证明对视角协变块(未来画面)施加一致性损失有确定性损害,提出只约束视角不变块(动作块、本体感知与价值)的 SCVC 方法;在 LIBERO-Plus 留出机位上闭环成功率较对照提升 12.2 个百分点,无需相机外参或测试时信息。

AI summary · AI generated

arXiv:2608.21402v1 Announce Type: new Abstract: World action models (WAMs) jointly denoise future video frames and robot actions, and the video prior is expected to generalize their control. Camera viewpoint change remains one of their hardest perturbation axes. We study a question specific to this model class: when training with same-state cross-view image pairs, on which output coordinates should a consistency loss be imposed?

跨视角一致性只约束动作块:世界动作模型对未见机位成功率提升 12 个百分点 | Embodied Intelligence Observer