Embodied Intelligence Observer

ForeTime-VLA:世界动作模型因果未来 token 蒸馏用于传送带操作

Original title: ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation

IndustryAI 66

Source: arXiv cs.ROPublish time unverified

arXiv:2608.20735v2 Announce Type: replace-cross Abstract: Manipulating moving objects requires a policy to anticipate contact events, yet vision-language-action (VLA) policies are commonly fine-tuned from the current observation alone. World action models (WAMs) learn predictive dynamics, but running a video-scale teacher or explicitly imagining future frames at deployment is costly.

ForeTime-VLA:世界动作模型因果未来 token 蒸馏用于传送带操作 | Embodied Intelligence Observer