Embodied Intelligence Observer

Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching

Research

Source: arXiv cs.ROPublish time unverified

arXiv:2602.00686v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable generalization capabilities in robotic manipulation tasks, yet their substantial computational overhead remains a critical obstacle to real-world deployment. Improving inference efficiency is therefore essential for practical robotic applications.

Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching | Embodied Intelligence Observer