Gemini 领跑,GPT表现平平,国立清华、英伟达提出双奖励孪生网络,告别算力焦虑。 作者丨张 璐 编辑丨幸丽娟 VLM大模型(VLM)在静态视觉与推理任务上已展现出强悍能力。然而在具身智能领域,当模型需要面对复杂的动态交互与机械臂控制时,它们能否准确评估智能体的动作与画面变化,依然是一个待检验的问题。 过去不少人以为,能“看懂画面”就意味着能“当好教练”。但对具身智能来说,从识别场景到精准评估动作进展,完全是两码事。 为了厘清各大模型在真实动作评估中的底细,国立清华大学与 NVIDIA 联合团队在最新研究《VLM-AR3L: Vision-Language Models for Absolute and Relative Rewards in Reinforcement Learning》中,把 Gemini 2.0、GPT-4.1 等主流大模型拉进同一个考场,对它们的视觉裁决能力做了一场综合评估。 论文链接:https://arxiv.org/pdf/2607.00483 01 考场设计:四大仿真套件、 10 大场景与“看图判进展”规则 研究团队选用的 10 个典型任务,涵盖了具…

9月2日,前字节跳动强化学习专家、前腾讯Robotics X智能体中⼼负责⼈孙鹏博⼠正式加入星尘智能
学术数据和真实数据完全是两回事;可解释性成落地“生死线”。 作者丨幸丽娟 编辑丨岑 峰 三年前,正值大模型爆发之年,麦肯锡就曾对全球交通与工业领域的从业者做了一项调研,问题很朴素:距离真正的智能化,还有多远? 答案出乎意料地悲观:还要10到20年。 而把时间线拉到现在,模型越来越大、智能体越来越多、时空数据挖掘的论文层出不穷。 现实依旧是,全球没有一座城市长期部署了基于强化学习的交通信号控制系统。那些在顶会上“超越所有SOTA”的预测模型,真正到了城市交通部门手里,他们的答案往往是:“这个方案,我不敢放手用。” 问题的症结在哪里? 在刚刚结束的 IJCAI 2026 Early Career Spotlight(早期职业焦点)演讲中,慕尼黑工业大学 (TUM) W2 教授黎子玥(Ziyue Li)给出了他的答案: 第一,学术研究长期停留在“干净数据”的舒适区,而公共部门面对的,永远是不完整、不规整、不可控的真实数据; 第二,可解释性不是论文里的加分项,而是从实验室走向城市落地的“生死线”。 黎子玥的履历横跨学界与工业界:香港科技大学工业工程与决策分析系博士,现任慕尼黑工业大学W2教授,…

作者| 宇航猿 编辑| 靖宇 很少有机器人,能让你第一眼就笑出来。 8 月 27 日,Hugging Face 旗下的 Pollen Robotics,开放了 Microduck 的预购。 一只 25 厘米高、不到 800 克重的双足机器鸭,四种配色,399 美元,圣诞节前发货。它能走路,能蹲下再站起来,被推倒了自己翻身,甚至能踩上一对可拆卸的轮滑鞋溜冰。嘴巴是铰接式的,低头叼起地上的袜子或马克笔,再直起身子,嘴里还叼着「战利品」。 Microduck 足球赛|图片来源:Hugging Face 它没有语音交互,不会说话,但有麦克风和扬声器,每一只 Microduck 在初始设置后会生成自己独有的「音色身份」。 官方的设定很明确—— 把它当一个活物,而不是一个助手。 如果你只看到了「可爱」,那你大概率低估了这只鸭子背后的东西。 01 拆箱就能玩的「电子宠物」 盒子里的东西很简单——机器人本体、一块电池、一根 USB-C 线、一只游戏手柄。充上电,配对手柄,鸭子就能动了。 出厂预装了 7 种行为策略,每一种都是经过强化学习训练的独立动作模型。用手柄推摇杆,Microduck 迈开两条小短…

谁不想在新年到来之际收到一只可爱的 AI 机器鸭呢? 距离 2026 年圣诞节还有四个月,在英伟达以约 129 亿美元收购消息曝光次日,AI 开源社区 Hugging Face 发布了一款名为 Microduck 的小机器人。它身高 25 厘米、重不足 800 克。演示视频里,它用喙叼着袜子摇摇晃晃地走路,被人推搡时会尽力稳住,即便倒下也能自己爬起来,甚至可以穿上轮滑鞋滑旱冰。 产品售价 399 美元(约合 2,700 元人民币),有奶油白、石墨灰、薰衣草紫、天空蓝四种配色可选,附赠一只游戏手柄。联合创始人克莱姆·德朗(Clem Delangue)在社交平台 X 上放出一段 50 秒的演示视频,几个小时内浏览量接近 40 万。 (来源:Pollen Robotics) 从语音互动,到“走”进真实世界 做开源大模型托管平台起家的 Hugging Face,入局具身智能的节点要追溯到 2024 年初,前特斯拉(Tesla)Optimus 团队科学家雷米·卡德内(Rémi Cadene)加入,主导开源机器人库 LeRobot 的开发;2025 年 4 月,Hugging Face 完成对法国…

arXiv:2507.00611v2 Announce Type: replace-cross Abstract: Preference-based Reinforcement Learning (PbRL) provides a promising alternative to heuristic reward design in complex robotic environments. However, PbRL often suffers from poor sample efficiency, requiring extensive and costly human feedback, which limits its real-world applicability. Prior work has proposed learning a reward model from demonstrations and fine-tuning it using preferences. However, when the model is a neural network, tran…
arXiv:2608.17512v2 Announce Type: replace Abstract: Although Large Vision-Language Models (VLMs) have significantly advanced embodied navigation, their direct deployment remains challenging, as existing methods often force VLMs into unnatural action spaces that misalign with their 2D pre-training priors, compounded by rigid reasoning schedules and inefficient memory management. To overcome these limitations, we propose TAMP-Nav, a unified framework for efficient embodied navigation. First, we in…
arXiv:2608.16153v4 Announce Type: replace Abstract: Robot manipulation requires policies that are both accurate and efficient, as robot control must respond to changing observations under tight latency constraints. Recent diffusion and flow policies are promising, but they often treat conditions as auxiliary signals rather than jointly evolving them with action trajectories. We find that this limitation can be effectively mitigated by a \textbf{simple yet effective unified condition-action model…
arXiv:2608.26571v1 Announce Type: cross Abstract: Contrastive reinforcement learning (CRL) scales effectively in goal-conditioned tasks by casting policy learning into a self-supervised contrastive objective. However, in a failure-terminated Markov decision process, established CRL considers pre-failure future goals only when constructing positive samples, without accounting for the probability mass removed by failure termination. Our theoretical analysis shows that this omission induces a syste…
arXiv:2608.27079v1 Announce Type: new Abstract: Pretrained vision-language-action (VLA) policies provide strong priors for robot manipulation, yet adapting them online to fine-grained biomedical tasks remains challenging. Task success often hinges on subtle, view-dependent visual cues, while task-level rewards provide little guidance about which regions matter, making it difficult to learn task-relevant visual grounding from limited real-robot interaction. Online adaptation is further constraine…
arXiv:2608.26800v1 Announce Type: new Abstract: We present an online learning framework that enables a bimanual robot to acquire diverse juggling patterns directly on physical hardware within minutes, even with a significant sim2real gap. One of the most important lessons from this work is that a model, even when far from reality, can be extremely useful for learning. This motivates a central philosophy of our approach: learning should build upon the robot's current knowledge rather than replace…
arXiv:2608.26739v1 Announce Type: new Abstract: Accurate trajectory tracking in cable-driven lower-limb rehabilitation robots is challenging because model uncertainty, external disturbances, joint constraints, and pull-only cable actuation can degrade nominal control performance. Conventional model-based controllers provide an interpretable control structure but remain sensitive to model mismatch, whereas fully learning-based control can reduce transparency and complicate constraint-aware operat…
强化学习之父Rich Sutton接受红杉访谈,把矛头对准当下最主流的大模型路线:互联网数据终有上限、模型上线后权重基本冻结,行业或已陷入“局部最优”。他创办Oak Lab,押注能从自身经验持续学习的Agent,重新定义AI演进路径。 说起强化学习之父 Rich Sutton,AI圈几乎没人陌生。 他写下的《苦涩的教训》,至今仍被很多人视为理解AI发展的重要框架:从长期看,能够随着计算规模扩展的通用方法,往往会战胜依赖大量人工知识设计的方案。 但最近接受红杉资本访谈时,Sutton却把矛头对准了今天最主流的大模型路线。在他看来,LLM当然是一次伟大的科学突破,却也可能成为《苦涩的教训》的反面案例。 原因很简单。今天的大模型虽然摆脱了大量人工规则,却重新被“人类已经知道的东西”限制住了。互联网数据终究有限,合成数据背后依然需要人来决定什么值得生成。 而更关键的是,模型仍然不具备人类那样的持续学习能力。 一个司机开了10年车,会从经验里形成大量直觉。但今天的大模型,大部分学习都集中在预训练和后训练阶段。模型上线后,即使每天和几百万人交互,核心权重依然基本被冻结。 在Sutton看来,今天的…

Clem Delangue, CEO of Hugging Face, said the Microduck is an “open-source robot you can teach new tricks with reinforcement learning.”
arXiv:2608.23831v2 Announce Type: replace Abstract: While reinforcement learning (RL) allows generalist robot policies to continually improve during deployment, the large model size of modern generalist policies, such as VLAs, poses a fundamental obstacle to effective RL improvement. In particular, their severe inference latency---which can lead to pauses or jerky movements---can alter the effective environment dynamics and, if not correctly accounted for, break the Markov assumption that RL rel…
arXiv:2606.00307v2 Announce Type: replace Abstract: Recent works have explored unifying SLAM with geometric foundation models (GFMs). However, directly using GFM predictions for tracking is highly sensitive to model capability and uncertainty, as geometric inaccuracies in the predictions can adversely affect pose estimation. To address this limitation, we present a decoupled framework that integrates classical feature-based SLAM with GFMs, which achieves higher quality and more consistent dense …
arXiv:2510.04724v2 Announce Type: replace Abstract: This paper introduces a methodology for task-specific design optimization of multirotor Micro Aerial Vehicles. By leveraging reinforcement learning, Bayesian optimization, and covariance matrix adaptation evolution strategy, we optimize aerial robot designs guided exclusively by their closed-loop performance in a considered task. Our approach systematically explores the design space of motor pose configurations while ensuring manufacturability …
arXiv:2608.25350v1 Announce Type: cross Abstract: Vision-language models (VLMs) have emerged as a powerful source of supervision for reinforcement learning, enabling agents to leverage rich semantic knowledge during training. Inspired by the success of preference-based reward learning (PbRL) in reinforcement learning from human feedback (RLHF), vision-language model generated image-based preferences provide an effective source for learning reward functions. This can be done by visually comparing…
arXiv:2608.25470v1 Announce Type: new Abstract: This work presents a transient heat-transfer model of an industrial automated tape laying (ATL) process designed to overcome the limitations of conventional thermal models in composite manufacturing. The model solves the heat-conduction equation with coupled advection, conduction, convection, and radiation. A key innovation is the implementation of an analytical view factor approach that accounts for finite emitter and tape widths, thereby correcti…
RoboHarness用编排异构策略,让VLA、WAM、RL和TAMP不再各说各话,并零样本完成长时序任务。 作者丨邓哲敏 编辑丨齐铖湧 想象这样一个任务:打开柜门,找到藏在里面的积木,然后把它们搭成一座桥。 对人来说,这是一件很简单的事,几个动作自然衔接,幼儿园小朋友也能顺利完成。 但对机器人来说,这个任务横跨多种完全不同的能力:找积木,需要视觉理解和语言指令跟随;开柜门,需要与环境进行复杂交互;搭积木,需要几何规划和精准操控并推测环境变化。 更麻烦的是,这些种能力分属多个不同“门派”的模型——VLA(视觉-语言-动作模型)、WAM(世界-动作模型)、RL 策略(强化学习)、TAMP(任务与运动规划)。它们各有所长,又彼此割裂,训练方式不同、输入格式不同、状态空间也不同。 今天在机器人领域,能力不缺,但协作缺位。 华为诺亚方舟实验室近期发布了一篇论文,提出了名为 RoboHarness 的系统,专门解决不同策略之间无法协作的问题。围绕这项工作,论文作者与我们(雷峰网)进行了交流。 论文:https://arxiv.org/abs/2607.18060 项目主页:https://www…
arXiv:2608.16153v3 Announce Type: replace Abstract: Robot manipulation requires policies that are both accurate and efficient, as robot control must respond to changing observations under tight latency constraints. Recent diffusion and flow policies are promising, but they often treat conditions as auxiliary signals rather than jointly evolving them with action trajectories. We find that this limitation can be effectively mitigated by a \textbf{simple yet effective unified condition-action model…
arXiv:2607.14393v2 Announce Type: replace Abstract: Human-in-the-loop Reinforcement Learning has become a popular approach for training, finetuning, and aligning robot behavior with user preferences. Our paper explores the feasibility of using brain signals via functional near-infrared spectroscopy (fNIRS) to modulate robot learning in simulation. We compare agents trained on passive (observational) versus active (demonstrative) interaction tasks, and test multiple methods for enhancing the RL a…
arXiv:2607.00160v2 Announce Type: replace Abstract: Modular reconfigurable robotic systems provide a scalable solution for cooperative surface operations in future lunar missions. However, cooperative cargo transportation remains challenging due to morphology-dependent topology changes, strong payload-induced coupling, long-horizon decision making, and safety constraints. This paper proposes a phase-decomposed reinforcement learning framework for cooperative cargo transport with distributed robo…
arXiv:2606.22027v3 Announce Type: replace Abstract: Reinforcement learning for robot manipulation is often bottlenecked by reward design, especially in long-horizon tasks: sparse success rewards provide weak supervision, while hand-crafted dense rewards are tedious to design and generalize poorly across tasks. Progress-based reward models offer a promising alternative by estimating how far an observation has advanced toward task completion, but existing approaches often require task-specific dem…
arXiv:2606.08015v3 Announce Type: replace Abstract: We propose Q-Guided Value-Gradient Matching (Q-VGM), an offline-to-online reinforcement learning (RL) method for fine-tuning flow-matching vision-language-action (VLA) policies with a learned Q-function. Classical off-policy actor-critic methods improve a policy by following the critic gradient $\nabla_A Q$, but applying this update to flow policies requires backpropagation through the multi-step denoising process (BPTT), which is costly and un…