Embodied Intelligence Observer

Mamba 选择性状态空间改进 SmolVLA 精度-复杂度权衡

Original title: Mamba-based Selective State Space Modeling Improves the Accuracy-Complexity Tradeoff of SmolVLA Vision-Language-Action Experts

IndustryAI 66

Source: arXiv cs.ROPublish time unverified

arXiv:2608.21407v1 Announce Type: new Abstract: Vision-language-action (VLA) models face a crucial tradeoff between their task success rate and the policy-call frequency. Executing a single action per inference ($N=1$) enables accurate robot control but comes at the cost of huge compute time overheads, making real-time implementation infeasible.

Mamba 选择性状态空间改进 SmolVLA 精度-复杂度权衡 | Embodied Intelligence Observer