REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs
Research
Source: arXiv cs.ROPublish time unverified
arXiv:2609.01215v1 Announce Type: cross Abstract: Most vision-language-action (VLA) models -- OpenVLA, $\pi_0$, RT-2, RDT-1B -- are monolithic: they emit raw motor commands or short action chunks without organizing behavior into reusable abstractions, so they degrade on long-horizon tasks and resist interpretation.