Embodied Intelligence Observer

UniMem:统一多模态记忆与控制的 VLA 模型

Original title: UniMem: Unifying Multimodal Memory and Control for Vision-Language-Action Models

IndustryAI 66

Source: arXiv cs.ROPublish time unverified

arXiv:2608.22869v1 Announce Type: new Abstract: While Vision-Language-Action (VLA) models have leveraged internet-scale pretraining and task-focused finetuning to achieve strong performance on long-horizon tasks, they often struggle with non-Markovian tasks that require memory. Existing approaches to memory typically involve additional Vision-Language-Models (VLMs) for long-term memory management, introducing a memory bottleneck and a fractured training pipeline.

UniMem:统一多模态记忆与控制的 VLA 模型 | Embodied Intelligence Observer