Act with Intent:为 VLA 模型蒸馏行为意图
Original title: Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models
IndustryAI 67
Source: arXiv cs.ROPublish time unverified
arXiv:2608.23478v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which motor command was demonstrated while leaving implicit the local objective served by the behavior under the instruction.