Embodied Intelligence Observer

策略感知模拟器学习的理论基础与高效算法

Original title: Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning

ResearchAI 65

Source: arXiv cs.LGPublish time unverified

arXiv:2605.29032v3 Announce Type: replace Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss. However, powerful RL optimizers inevitably exploit minor model inaccuracies, leading to simulator exploitation and a reality gap where policies succeed in simulation but fail in the real world.

策略感知模拟器学习的理论基础与高效算法 | Embodied Intelligence Observer