Mamba 选择性状态空间改进 SmolVLA 精度-复杂度权衡
Original title: Mamba-based Selective State Space Modeling Improves the Accuracy-Complexity Tradeoff of SmolVLA Vision-Language-Action Experts
IndustryAI 66
Source: arXiv cs.ROPublish time unverified
arXiv:2608.21407v1 Announce Type: new Abstract: Vision-language-action (VLA) models face a crucial tradeoff between their task success rate and the policy-call frequency. Executing a single action per inference ($N=1$) enables accurate robot control but comes at the cost of huge compute time overheads, making real-time implementation infeasible.