Beijing Humanoid Robotics Innovation Center sweeps the World Humanoid Robot Games while CEO Xiong Youjun argues the industry has reached embodied AI's pre-ChatGPT moment.

As China builds embodied-intelligence training grounds nationwide, a more than 99% shortfall in physical-interaction data risks keeping humanoid robots stuck at the demo stage.

如果说,基础模型是具身大脑的“大学”通识课;那么,垂域模型则是机器人在工业领域的“专业课”。 作者丨董子博 编辑丨林觉民 今天要做世界模型,所有人都知道要做预训练,来让具身大脑变得更聪明;也知道要做后训练,让机器人可以在专门的任务下得到锻炼、做得更好。 和不少公司聊过,AI 科技评论发现,谈到中训练(Mid-train)的公司却很少。 而实际上,在具身的实际落地中,却存在一些原有方案不易解决的问题:当预训练的 VLA 模型,没法在具体的场景中发挥全部智能;而后训练却有时间、成本的门槛;在训练技能的环节,也容易出现数采、训练和调试的重复工作——VLA 落地要真正提效,中训练或许才是破局的良方。 AI科技评论独家获悉,乐聚机器人将推出面向工业场景构建的“垂域具身智能模型”——KUAVO VLA。据了解,该模型以通用VLA 基础模型为能力起点,使用 600+ 小时KUAVO 同构型真机数据进行了中训练,将本体适配、基础操作和工业场景的共性能力沉淀到模型中。 乐聚机器人 KUAVO VLA 赋能下,在不少场景中可以实现大约一倍的提效,在某些任务中甚至可以做到100% 的成功率。 模型中训练,会…

arXiv:2608.25940v1 Announce Type: new Abstract: Physical AI models are evaluated on suites of benchmarks that differ across model reports, leaving the model-by-benchmark matrix sparse and the relationship between benchmarks unmeasured. We construct a matrix of 51 models on 12 physical AI benchmarks, selected from a registry of 51 benchmarks and 152 models by reporting density, combining scores from model cards and benchmark papers with our own evaluation runs under each benchmark's official prot…
Insider Brief Qualcomm Technologies plans to launch a long-term investment initiative in Japan focused on robotics, physical AI and industrial automation, including a new research and development hub intended to connect companies, startups and universities. Qualcomm said the initiative is designed to support robotics development from early-stage technology work through commercial deployment and expansion into […]
Amazon Web Services (AWS), an Amazon.com, Inc. company (NASDAQ: AMZN), and NVIDIA (NASDAQ: NVDA) today announced a major expansion of their strategic collaboration to meet surging global demand for AI infrastructure as demand continues to accelerate.
物理AI的赛点时刻,是建立人机协同的信任。 作者丨倪 萍 编辑丨王瑞昊 过去几天,北京亦庄经历了一场机器人的狂欢。 约6万平方米的展馆内,上千台形态各异的机器人忙着做菜、打球、递咖啡、分拣快递,甚至还有专门铲猫砂、处理猫毛的“机械铲屎官”。从工厂到家庭,从生产到生活,硅基生命体们在每一个细分场景中卖力展示着解决现实问题的能力。 由此可见,具身智能的叙事焦点,正从“表演特技”加速转向“实景作业”。2026世界机器人大会首次设立采购日,专业买家携需求清单进场,供需直接对接,现场销售额累计突破2亿元。 黄仁勋所描绘的物理AI世界正快速到来:AI跨出屏幕边界,通过各式各样的物理“身体”,直接参与并改变人类生活。 这意味着更快的交付速度、更广泛的场景覆盖,以及更低成本的商业闭环。 然而,眼下的人形机器人、轮式机器人跨过量产门槛仍需时日。有能力率先实现人机深度协作图景的,很可能是AI汽车。 在2026世界机器人大会火山引擎具身智能技术与应用论坛上,AIVA总裁、产品经理李博给出了一个判断:物理AI这条赛道,最终比拼的或许不只是模型有多强大、硬件有多灵巧,而是它能否与人之间,构建起一段真实可感的信任…
RoboHarness用编排异构策略,让VLA、WAM、RL和TAMP不再各说各话,并零样本完成长时序任务。 作者丨邓哲敏 编辑丨齐铖湧 想象这样一个任务:打开柜门,找到藏在里面的积木,然后把它们搭成一座桥。 对人来说,这是一件很简单的事,几个动作自然衔接,幼儿园小朋友也能顺利完成。 但对机器人来说,这个任务横跨多种完全不同的能力:找积木,需要视觉理解和语言指令跟随;开柜门,需要与环境进行复杂交互;搭积木,需要几何规划和精准操控并推测环境变化。 更麻烦的是,这些种能力分属多个不同“门派”的模型——VLA(视觉-语言-动作模型)、WAM(世界-动作模型)、RL 策略(强化学习)、TAMP(任务与运动规划)。它们各有所长,又彼此割裂,训练方式不同、输入格式不同、状态空间也不同。 今天在机器人领域,能力不缺,但协作缺位。 华为诺亚方舟实验室近期发布了一篇论文,提出了名为 RoboHarness 的系统,专门解决不同策略之间无法协作的问题。围绕这项工作,论文作者与我们(雷峰网)进行了交流。 论文:https://arxiv.org/abs/2607.18060 项目主页:https://www…
将动作从输出端搬到调度端,推理速度提升1.79倍。 作者丨邓哲敏 编辑丨齐铖湧 机器人想要真正走进千行百业、千家万户,有一道绕不开的坎:它得足够快。 不说快到能打羽毛球那,至少得快到让人觉得这玩意儿能用。一个机械臂每动一下都要停顿两三秒,哪怕成功率再高,在工业场景里算下来的投入产出比也很难说服人。 VLA 模型是当前具身智能领域最受关注的技术路线之一。它把视觉感知、语言理解和动作生成整合进一套模型,让机器人能够依据语言指令完成复杂操作任务。但 VLA 模型有一个部署层面的老大难问题:推理太慢。 每个控制时刻,模型都要把一个庞大的视觉语言骨干网络跑一遍,计算量极大,导致延迟居高不下,实时控制频率难以保证。 这个问题并不新鲜,学界已经有不少人在尝试解决。今年 IJCAI 收录的一篇论文认为,现有的大多数方案都走偏了一步。AI科技评论(雷峰网公众号)联系到一作于闻达,围绕高效 VLA 方法进行交流。 论文原文:https://arxiv.org/abs/2601.19634 01 大家都在剪视觉, 但问题真的在这儿吗 现有的高效 VLA 方法,主流路线是在视觉侧做文章,剪掉“不重要”的视觉 …
2026 WRC 现场,三台机器人默契配合,正完成工业制造拆垛、分拣、搬运的全流程工作;几步之外,观众在零售舱前下单,等机器人递来饮料;厨房里,一顿饭从食材处理一路做到最后的清洁;在生产线之外的赛博舞台区域,NAVIAI i3 单膝跪地献花,WA2 弯腰接过。 这些画面,来自浙江人形机器人创新中心有限公司的展台,在这几块相邻的区域中,浙江人形把机器人的工作半径从生产线拉到人的日常,机器人的每一条动作轨迹、每一次执行,都在诠释浙江人形展台主题“Hello, World!具身在场,让向往真正发生”。 但如果只把这次展出读成“场景丰富”,其实有点可惜。 把几个画面连起来看,真正被拉长的,是机器人承担一件事的长度:一笔订单要从下单走到交付,一顿饭要从备餐做到收尾,一段工序要接着上一段继续往下跑。 “向往”也因此变得具体,机器多接住一些事情,劳动就多一种选择,生活多一分从容。 只是,场景从来不是终点。 当“进场景”已经成为行业的共同语言,差距也开始出现在后半程:机器人不仅要进得去,还得证明自己能在那里留下来。 机器人开始干活 所有人都成了“监工” 同一个 WRC,不同的人有不同的算盘。 投资人算…
As China builds scores of embodied-intelligence data centers, this year's World Robot Conference showed an industry reorienting from building robot bodies to harvesting the data that trains their brains.
The $200 million extension comes just months after the physical AI startup reached a $2 billion valuation.
Humanoid robots are having a moment in China. The popular machines are part of the country’s strategy to bring artificial intelligence into daily life. Embedding the technology into physical systems—an idea called embodied AI—was a key facet of China’s latest five-year plan, and companies here are already world leaders in humanoids. Nearly 90% of the…

The 11th World Robot Conference in Beijing Yizhuang saw hundreds of companies demonstrate robots doing real work, even as closed-door sessions debated a divergent technical frontier for embodied intelligence.

告别理想仿真,中国学术团队用物理工程方案回应落地挑战。 作者丨张璐 编辑丨岑峰 过去一年多里,具身智能似乎按下了快进键。 从 Physical Intelligence 发布的π₀到开源社区追捧的OpenVLA,VLA大模型一路高歌猛进。在各种标准基准测试和精心搭建的实验室 Demo 中,机械臂叠衣服、拿取物品流畅得宛如人类,动辄刷出 90% 甚至 95% 以上的成功率。 然而,当全行业都在欢呼“机器人的 GPT 时刻”到来时,前沿学术界与产业界却爆发了一场剧烈的反思:这些动辄数十亿参数的具身大模型,真的学会“举一反三”了吗?还是在“死记硬背”? 在主流学术视角中,VLA 大模型像是给机器人装上了具备常识与规划能力的“大脑”。但要让机械体真正干好精细活,仅有一个聪明的大脑远远不够。 最新实验数据表明,面对物品位置的随机挪动或新指令组合时,纯端到端大模型往往会受到“空间过拟合”的约束,末端执行成功率会出现剧烈衰减。 这恰恰说明:具身智能要真正走出实验室,不仅需要大模型提供常识和高层决策,更需要扎实的工程技术去打造精准、稳健的“小脑”。 下面雷峰网盘点入选 IJCAI 2026 的 6 项…

当机器人行业的大量叙事还停留在融资轮次与概念发布时,越疆科技已经交出了一份由产线验收单写成的中期答卷。 8月24日,越疆科技发布2026年上半年中期业绩。报告期内,公司实现营业收入约人民币3.2亿元,同比增长106.6%,经营规模实现翻倍增长;毛利约人民币1.5亿元,同比增长约100%,毛利率稳定维持在47.4%。研发投入约人民币1.0亿元,同比增长148.4%。具身智能机器人业务收入突破人民币4,500万元,同比增长超20倍,占营收比例提升至14.3%。 在工业机器人行业产量同比增长28%的产业背景下,越疆以远超行业均速的表现,印证了“协作机器人+具身智能”双轮驱动战略的强劲动能。 一、深度锚定工业主场,协作机器人筑牢全球龙头地位 2026年上半年,工业制造领域收入约人民币2.0亿元,同比增长142.4%,占产品销售收入的62.8%,为公司第一大收入来源。协作机器人实现收入约人民币2.4亿元,同比增长79.1%。 根据灼识咨询报告,按2025年销量统计,越疆协作机器人全球市占率达13.2%,位列全球第一。全球累计出货量突破10万台,业务落地100余个国家和地区,长期服务80余家世界5…

HaReCAP 为递归 LLM 智能体补充离线编译的叶级反射规则,仅在规则能唯一确定合法动作时跳过 LLM 调用,在 Robotouille 与 ALFWorld 上将 token 消耗降低 14.7%-20.1%,且不改变原递归控制流。AI summary
arXiv:2605.27533v2 Announce Type: replace Abstract: Periods of heightened arousal or restlessness can interfere with children's ability to focus, self-regulation, and physically calm. Technologies that encourage embodied self-regulation through tactile interaction may provide a simple and accessible means of promoting calmness. This paper investigates how interaction with a pocket-sized tactile device influences physiological and behavioral markers of calmness in typically developing children. B…
arXiv:2603.16673v5 Announce Type: replace Abstract: Embodied robotic systems increasingly rely on large language model (LLM)-based agents to support high-level reasoning, planning, and decision-making during interactions with the environment. However, invoking LLM reasoning introduces substantial computational latency and resource overhead, which can interrupt action execution and reduce system reliability. Excessive reasoning may delay actions, while insufficient reasoning often leads to incorr…
arXiv:2509.23155v3 Announce Type: replace Abstract: Robotic manipulation benefits from foundation models that describe goals, but today's agents still lack a principled way to learn from their own mistakes. We ask whether natural language can serve as feedback, an error-reasoning signal that helps embodied agents diagnose what went wrong and correct course. We introduce LaGEA (Language Guided Embodied Agents), a framework that turns episodic, schema-constrained reflections from a vision language…
arXiv:2608.21928v1 Announce Type: cross Abstract: In embodied AI, safety risk can be latent: a benign instruction and a safe scene become hazardous only when composed. Prior work has advanced embodied safety by varying visual contexts or evaluating execution-time dynamics, but the complementary axis of fixing the scene and varying only the instruction remains underexplored. We introduce GuardianBench, an instruction-contrastive benchmark grounded in international safety standards that isolates t…
arXiv:2608.23354v1 Announce Type: new Abstract: Autonomous indoor navigation requires both semantic understanding and precise geometric control. We propose OptiSight, a hybrid framework that combines Vision-Language Model reasoning with deterministic visual servoing through a finite-state Chain-of-Thought architecture. Grounded-SAM localizes open-vocabulary targets, while camera projection geometry converts visual observations into navigation commands without requiring dense mapping. The VLM is …
arXiv:2608.23138v1 Announce Type: new Abstract: Vision-language-action (VLA) models often expose spatial grounding through autoregressive text coordinates or opaque action tokens, creating brittle interfaces between multimodal reasoning and robot execution. We present Pointing-VLA, a typed hidden-state spatial readout built on Embodied-R1. Geometry-specific heads predict normalized points, object-functional grounding (OFG) heatmaps, and visual trajectories without serializing geometry as text. F…
arXiv:2608.23000v1 Announce Type: new Abstract: Fully online embodied learning requires synaptic adaptation to acquire new behaviors while preserving previously learned dynamics during ongoing interaction. We extend the Predictive-Coding-inspired Variational Recurrent Neural Network (PV-RNN) to continuously adapt its synaptic weights and propose Free-Energy-Gated Plasticity (FEGP), which regulates the effective learning rate according to variational free energy. In real-time physical human--robo…
arXiv:2608.21899v1 Announce Type: new Abstract: Human-in-the-loop real-world reinforcement learning enables rapid acquisition of effective robotic manipulation policies for individual tasks, often within tens of minutes. Yet it remains unclear how to extend this paradigm to continual learning, where a single policy must acquire new skills without losing previously learned behaviors. Existing real-world continual learning methods do not explicitly constrain prior behaviors, leading to severe cata…
GOLEM 是基于 Unitree H1-2 的开源电动车电池拆卸系统架构,行走、操作、导航、空间记忆等模块解耦并配 MuJoCo/IsaacLab 数字孪生;实测从现代 Ioniq 5 电池包抓取松动紧固件,系留状态成功率 97%、自由站立降至 87%。AI summary
arXiv:2608.21416v1 Announce Type: new Abstract: Embodied artificial intelligence (AI) must be tested in the clinical environments where it will operate, but building realistic, robot-testable settings is costly and difficult to scale. Here we show that routine clinic images can be transformed into operational digital twins for task-based evaluation of embodied AI. Using 39 ophthalmic clinic scenes, we converted single photographs into editable, simulator-ready environments and assessed reconstru…