Dream2Reward: Transition-Alignment Reward Models from Positive Demonstrations for Robotic Manipulation
Research
Source: arXiv cs.ROPublish time unverified
arXiv:2608.18787v1 Announce Type: new Abstract: Learning robotic policies requires dense rewards that remain informative when behavior departs from successful demonstrations. Progress-based rewards estimate how far an observation has advanced along a nominal successful trajectory, but may remain high after an incorrect transition. We introduce Dream2Reward, which learns a language-conditioned successful latent transition field from positive demonstrations.