Embodied Intelligence Observer

AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action Models

Research

Source: arXiv cs.ROPublish time unverified

arXiv:2608.29208v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models, built upon Vision-Language Models (VLMs), have significantly enhanced robotic capabilities by leveraging internet-scale knowledge and multimodal reasoning. However, the intensive computational overhead of VLAs constrains on-device deployment, hindering real-time responses to environmental changes.

AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action Models | Embodied Intelligence Observer