具身智能观察

AcrossVAM1.0:文本辅助机器人视频预测的粒子世界建模

原标题:AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction

技术动态AI 70

来源:arXiv cs.RO发布时间待核实

arXiv:2608.28491v1 Announce Type: cross Abstract: Predicting robot videos requires both precise motion reasoning and preservation of high-frequency appearance, yet monolithic pixel models entangle these objectives and often conceal their progress behind a strong last-frame baseline. We present AcrossVAM1.0, a lightweight, text-assisted video action model that factorizes future prediction into object-centric motion and dense appearance.

AcrossVAM1.0:文本辅助机器人视频预测的粒子世界建模 | 具身智能观察