Embodied Intelligence Observer

Guardian: Detecting Robotic Planning and Execution Errors with Vision-Language Models

Industry

Source: arXiv cs.ROPublish time unverified

arXiv:2512.01946v4 Announce Type: replace Abstract: Robust robotic manipulation requires reliable failure detection and recovery. Although recent Vision-Language Models (VLMs) show promise in robot failure detection, their generalization is severely limited by the scarcity and narrow coverage of failure data. To address this bottleneck, we propose an automatic framework for generating diverse robotic planning and execution failures across both simulated and real-world environments.