具身智能观察

策略感知模拟器学习的理论基础与高效算法

原标题:Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning

技术动态AI 65

来源:arXiv cs.LG发布时间待核实

arXiv:2605.29032v3 Announce Type: replace Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss. However, powerful RL optimizers inevitably exploit minor model inaccuracies, leading to simulator exploitation and a reality gap where policies succeed in simulation but fail in the real world. We propose that the objective for learning simulators should be strategic robustness rather than predictive accuracy, and formulate this as a zer…