从一比特失败信号扩展安全目标条件策略学习
原标题:Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals
产业动态AI 68
来源:arXiv cs.RO发布时间待核实
arXiv:2608.26571v1 Announce Type: cross Abstract: Contrastive reinforcement learning (CRL) scales effectively in goal-conditioned tasks by casting policy learning into a self-supervised contrastive objective. However, in a failure-terminated Markov decision process, established CRL considers pre-failure future goals only when constructing positive samples, without accounting for the probability mass removed by failure termination.