LD4WAM:从人类视频学习隐动力学供世界动作模型
Original title: LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models
IndustryAI 69
Source: arXiv cs.ROPublish time unverified
arXiv:2608.22403v1 Announce Type: new Abstract: Human video is playing an increasingly central role in training World Action Models (WAMs), owing to its diversity and low collection cost relative to teleoperated robot data. However, most WAMs learn from such video only by predicting pixel-level future frames, giving dynamics that are not directly actionable, whereas motion retargeting recovers directly actionable actions but leaves a large visual gap across embodiments.