Embodied Intelligence Observer

从一比特失败信号扩展安全目标条件策略学习

Original title: Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals

IndustryAI 68

Source: arXiv cs.ROPublish time unverified

arXiv:2608.26571v1 Announce Type: cross Abstract: Contrastive reinforcement learning (CRL) scales effectively in goal-conditioned tasks by casting policy learning into a self-supervised contrastive objective. However, in a failure-terminated Markov decision process, established CRL considers pre-failure future goals only when constructing positive samples, without accounting for the probability mass removed by failure termination.

从一比特失败信号扩展安全目标条件策略学习 | Embodied Intelligence Observer