极简模态掩码即可强化双臂 VLA 鲁棒性
Original title: Robust Bimanual Vision-Language-Action Models via Embarrassingly Simple Modality Masking
IndustryAI 65
Source: arXiv cs.ROPublish time unverified
arXiv:2608.22419v1 Announce Type: new Abstract: Query-based Vision-Language-Action (VLA) models offer low-latency inference that is attractive for bimanual robotic manipulation, but we observe that they can still exhibit discontinuous actions and execution failures in complex dual-arm tasks. We hypothesize that unstable multi-view and language fusion is one contributing factor in these failures, often coinciding with attention spreading to distracting regions.