CounterAlign: Counterfactual Supervision for Vision-Language-Action Models
Industry
Source: arXiv cs.ROPublish time unverified
arXiv:2608.21740v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are typically trained with behavior cloning (BC) on expert demonstrations. However, BC provides only positive supervision for expert actions, without explicit negative supervision indicating which actions are instruction-inconsistent or otherwise inappropriate.