EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents
技术动态
来源:arXiv cs.RO发布时间待核实
arXiv:2609.01281v1 Announce Type: new Abstract: Vision-language-action (VLA) models map visual observations and language instructions directly to robot actions, but long-horizon tasks require more than action prediction. An agent must coordinate perception, planning, execution, progress verification, and recovery as the physical state evolves.