Embodied Intelligence Observer

TemporalFlow-VLA:长时程操作的物理接地执行历史学习

Original title: TemporalFlow-VLA: Learning Physically Grounded Execution History for Long-Horizon Robot Manipulation

IndustryAI 80

Source: arXiv cs.ROPublish time unverified

arXiv:2608.26821v1 Announce Type: new Abstract: Vision-language-action (VLA) models leverage pretrained vision-language representations for robot control, yet simply adding historical frames does not reliably capture recent physical change. This is especially problematic in multi-stage manipulation, where visually similar states may require different actions depending on prior execution.

TemporalFlow-VLA:长时程操作的物理接地执行历史学习 | Embodied Intelligence Observer