具身智能观察

从一比特失败信号扩展安全目标条件策略学习

原标题:Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals

产业动态AI 68

来源:arXiv cs.RO发布时间待核实

arXiv:2608.26571v1 Announce Type: cross Abstract: Contrastive reinforcement learning (CRL) scales effectively in goal-conditioned tasks by casting policy learning into a self-supervised contrastive objective. However, in a failure-terminated Markov decision process, established CRL considers pre-failure future goals only when constructing positive samples, without accounting for the probability mass removed by failure termination.

从一比特失败信号扩展安全目标条件策略学习 | 具身智能观察