具身智能观察

Q-VGM: Q-Value-Gradient Matching for Offline-to-Online Reinforcement Learning of Flow-Matching VLA

产业动态

来源:arXiv cs.RO发布时间待核实

arXiv:2606.08015v3 Announce Type: replace Abstract: We propose Q-Guided Value-Gradient Matching (Q-VGM), an offline-to-online reinforcement learning (RL) method for fine-tuning flow-matching vision-language-action (VLA) policies with a learned Q-function.