AcrossVAM1.0:文本辅助机器人视频预测的粒子世界建模
原标题:AcrossVAM1.0: Particle World Modeling for Text-Assisted Robot Video Prediction
技术动态AI 70
来源:arXiv cs.RO发布时间待核实
arXiv:2608.28491v1 Announce Type: cross Abstract: Predicting robot videos requires both precise motion reasoning and preservation of high-frequency appearance, yet monolithic pixel models entangle these objectives and often conceal their progress behind a strong last-frame baseline. We present AcrossVAM1.0, a lightweight, text-assisted video action model that factorizes future prediction into object-centric motion and dense appearance.