arXiv:2509.23155v3 Announce Type: replace Abstract: Robotic manipulation benefits from foundation models that describe goals, but today's agents still lack a principled way to learn from their own mistakes. We ask whether natural language can serve as feedback, an error-reasoning signal that helps embodied agents diagnose what went wrong and correct course. We introduce LaGEA (Language Guided Embodied Agents), a framework that turns episodic, schema-constrained reflections from a vision language…
arXiv:2608.22731v1 Announce Type: cross Abstract: Nonverbal behavior generation systems for virtual agents often take an utterance as input and generate nonverbal behaviors that emphasize or illustrate the content of the verbal channel. However, human nonverbal behavior is shaped by more than the content of the speech. It is also influenced by speaker roles, interpersonal relationships, social context, and the cognitive and emotional states of the interactants. As a result, the nonverbal channel…
arXiv:2608.21628v1 Announce Type: cross Abstract: Black-box VR and 3D applications are difficult to regression test because observable failures depend on where a tester moves, what objects are visible, and which views are captured. Manual exploratory testing can find such failures, but its evidence is time-consuming to reproduce; systematic sweeps are reproducible, but they lack semantic guidance and spend exploration budget on low-value viewpoints. We observe that an LLM can make the high-level…
arXiv:2608.23478v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which motor command was demonstrated while leaving implicit the local objective served by the behavior under the instruction. Future-based supervision enriches action learning with frames, latent observations, trajectories, or motion representations, but these signals capture pa…
arXiv:2608.23354v1 Announce Type: new Abstract: Autonomous indoor navigation requires both semantic understanding and precise geometric control. We propose OptiSight, a hybrid framework that combines Vision-Language Model reasoning with deterministic visual servoing through a finite-state Chain-of-Thought architecture. Grounded-SAM localizes open-vocabulary targets, while camera projection geometry converts visual observations into navigation commands without requiring dense mapping. The VLM is …
arXiv:2608.23320v1 Announce Type: new Abstract: Industrial demand changes the paradigms of production. Due to smaller batch sizes and more variations in products, companies face a growing challenge to adopt more adaptive production systems. In particular, robot-based automation is usually static and fails to respond to constantly changing processes. Vision-Language-Action (VLA) Models are a promising opportunity to mitigate this challenge by generating robot actions based on the observed system …
arXiv:2608.23224v1 Announce Type: new Abstract: Retrieval can efficiently and effectively augment a frozen vision--language--action (VLA) policy without retraining, yet retrieved text becomes a control intervention once it enters the executed prompt. In a matched audit, raw appended text reduces mean success from 92.47\% to 3.00\%, while meaningful and length-matched meaningless appends both fail on all 500 states. This result identifies \emph{prompt-form collapse}: changing the instruction form…
arXiv:2608.23138v1 Announce Type: new Abstract: Vision-language-action (VLA) models often expose spatial grounding through autoregressive text coordinates or opaque action tokens, creating brittle interfaces between multimodal reasoning and robot execution. We present Pointing-VLA, a typed hidden-state spatial readout built on Embodied-R1. Geometry-specific heads predict normalized points, object-functional grounding (OFG) heatmaps, and visual trajectories without serializing geometry as text. F…
arXiv:2608.22990v1 Announce Type: new Abstract: Vision-language-action (VLA) models have made general-purpose robot manipulation increasingly plausible by conditioning robot actions on natural-language instructions. A key test of such generality is whether policies actually follow language instructions. Yet many manipulation benchmarks leave this ability underdetermined: the intended object or destination is often visually salient or uniquely feasible, allowing policies to succeed without ground…
arXiv:2608.22896v1 Announce Type: new Abstract: Robotic navigation in human environments requires a spatio-temporal semantic representation that can rec- oncile open-vocabulary perception with long-term environmental changes. While foundation models provide strong zero-shot recognition, their predictions are intermittent and view-dependent, and naively integrating them into mapping pipelines leads to identity drift and stale semantics over time. We present SuperMap, a 4D spatio-temporal mapping …
arXiv:2608.22869v1 Announce Type: new Abstract: While Vision-Language-Action (VLA) models have leveraged internet-scale pretraining and task-focused finetuning to achieve strong performance on long-horizon tasks, they often struggle with non-Markovian tasks that require memory. Existing approaches to memory typically involve additional Vision-Language-Models (VLMs) for long-term memory management, introducing a memory bottleneck and a fractured training pipeline. Conditioning on multiple histori…
arXiv:2608.22800v1 Announce Type: new Abstract: Ensuring reliability in uncertain environments remains difficult for long-horizon robotic manipulation. End-to-end VLA models are data-heavy and opaque, making diagnosis and verification difficult. Hierarchical pipelines are more interpretable, but their plans are often weakly grounded in observations, weakly aligned with low-level actions, and computed without online feedback, leading to open-loop behavior and hallucinations. To address these issu…
arXiv:2608.22449v1 Announce Type: new Abstract: Forecasting dexterous hand motions from egocentric observations is fundamental to intelligent interactive systems. Existing VLM-based methods typically map observations directly to future motions, overlooking the underlying manipulation process that governs hand-object interactions. Moreover, end-to-end optimization couples manipulation learning with motion synthesis, causing motion-generation gradients to interfere with the pre-learned manipulatio…
arXiv:2608.22419v1 Announce Type: new Abstract: Query-based Vision-Language-Action (VLA) models offer low-latency inference that is attractive for bimanual robotic manipulation, but we observe that they can still exhibit discontinuous actions and execution failures in complex dual-arm tasks. We hypothesize that unstable multi-view and language fusion is one contributing factor in these failures, often coinciding with attention spreading to distracting regions. To improve robustness, we introduce…
arXiv:2608.22149v1 Announce Type: new Abstract: LLMs generate fluent plans for robots but routinely violate the syntactic and se8mantic constraints they must satisfy to execute, and existing remedies trade formal guarantees against plan quality: soft methods (affordance scoring, grounded decoding) give no guarantee, while symbolic planners (LLM+P) discard the LM's commonsense. We propose \textbf{Meta-Ctrl}, a constrained-decoding framework that guarantees the encoded constraints while preserving…
arXiv:2608.22035v1 Announce Type: new Abstract: Robot foundation models have substantially advanced perception and control, but natural human-robot collaboration requires more than executing isolated commands. A robot must recognize ambiguity, maintain context across turns, communicate its intentions, and revise ongoing behavior as the user's intent changes. We present $\scriptstyle\mathsf{Ludi}_{\scriptscriptstyle 0.1}$, an agentic system for socially intelligent robots that integrates interact…
arXiv:2608.21740v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are typically trained with behavior cloning (BC) on expert demonstrations. However, BC provides only positive supervision for expert actions, without explicit negative supervision indicating which actions are instruction-inconsistent or otherwise inappropriate. Reinforcement learning (RL) can provide such corrective signals, but often relies on externally specified rewards or curated non-expert data, both of whic…
arXiv:2608.21440v1 Announce Type: new Abstract: Vision-language-action (VLA) models have advanced end-to-end autonomous driving by leveraging foundation models for semantic reasoning and long-tail generalization. However, their planning performance remains limited in complex driving environments because image-only representations inadequately capture planning-relevant road geometry and topology. In this paper, we propose Geo-VLA, a plug-and-play framework that enhances VLA models by learning geo…
对比人类与 LLM 生成主题分析在涉及弱势群体的人机交互研究中的表现,发现两者主题存在系统性分歧,LLM 分析有边缘化、误代表参与者经历的风险,建议敏感研究场景审慎使用并保留人工复核。AI summary
arXiv:2608.21407v1 Announce Type: new Abstract: Vision-language-action (VLA) models face a crucial tradeoff between their task success rate and the policy-call frequency. Executing a single action per inference ($N=1$) enables accurate robot control but comes at the cost of huge compute time overheads, making real-time implementation infeasible. On the other hand, executing longer action horizons before replanning ($N\gg1$) reduces compute complexity, but inevitably degrades the system's success…
arXiv:2608.21387v1 Announce Type: new Abstract: Service robots are increasingly deployed in elderly-care facilities to alleviate caregiver workload and enhance the quality of daily care. However, most existing studies focus on isolated service functions and lack integrated capabilities for continuous companionship, natural interaction, and safety monitoring. In this paper, we present an intelligent companion robot system that unifies active visual human-following, real-time LLM-driven speech int…
General Intuition, the startup building a foundation model that trains generalized AI agents how to move through space and time, is in talks to raise at a $6 billion pre-money valuation from new investors including Valor Ventures, Point72 Ventures, and Seven Seven Six.
8 月 19 日下午,世界机器人大会分论坛“具身智能大模型的进化之路:从技术分级到产业落地”在北人亦创国际会展中心举行。生数科技创始人兼首席科学家、ACM/IEEE/AAAI Fellow 朱军发表主旨演讲,发布团队最新研究成果,并系统阐述通用世界模型的五级发展路线。 朱军表示:“从大模型发展的视角来看,我们希望构建的,不只是服务于某个任务或场景的专用模型,而是一个能够理解世界、预测未来并采取行动的通用基座模型。” 从第一性原理出发,定义通用世界模型 世界模型已经延伸到视频生成、环境模拟、机器人决策和动作控制等方向,但行业对“什么样的世界模型才称得上通用”仍缺少共识。 人在学习骑自行车或开车时,会从最初的不协调逐渐变得稳定、准确。其重要原因,是大脑在与外部世界持续互动的过程中,形成了一个关于世界如何运行的“内部模型”。 通用世界模型同样需要三项彼此关联的能力: “通用世界模型不是单一的生成器、模拟器、机器人动作模型或策略模型,也不是这些能力的割裂片段或简单线性串联,而是三种能力耦合在一起的闭环反馈系统。”朱军强调,行动不仅是模型的输出,还会改变环境、产生新信息,并进入下一轮理解、预测和…

作者|李苏 编辑| 郑玄 一场特殊的网球比赛,一次数字智能向物理智能的跨越。 人形机器人挥出第一拍之前,没有人确定它能不能接住对面飞来的球。 事实上,球贴着边线飞来,机器人横移数步挥拍回击,球落在界内。它接住了,比赛继续。 这是 8 月 22 日第二届世界人形机器人运动会上的真实场景。开幕式网球表演赛上,一台人形机器人在全球直播镜头下,与网坛顶尖运动员展开全自主对抗。 发球、正手、反手、接发、底线调动、网前截击,整套动作连贯流畅。双打环节,它与人类队友实时协作,根据场上局势调整战术。高速攻防中,它跑过大半个球场完成极限救球,摔倒后还能自主站起。 这是全球首次人与机器人的大型赛事直播,做出这款机器人的公司叫做银河通用——这也其被定义为 AstraTennis 时刻。 很多人恍惚想起十年前,AlphaGo 在棋盘上战胜李世石,彼时,人工智能在数字世界的复杂博弈中抵达顶峰。但那时候,还仅仅是人脑和机器之间智力的对决。 堪堪十年光景之后,机器人在另一个真实的赛场上接住了世界名将的回球——数字智能开始向物理智能延伸。 支撑这场表演赛的,是银河通用自主研发的具身智能大模型银河星脑,以及核心技术平台…

北京,2026年8月19日—23日,2026世界机器人大会(WRC)在北京举行。 本届大会,智身科技以"真正干活的机器人"为核心定位,通过五大展区集中展示从核心关节、整机本体、机械臂到运动控制(小脑)、智能决策(大脑)的全栈技术能力,回答行业之问:机器人什么时候能真正创造社会价值? 01.展台即现场:从零件到整机,从表演到干活 走进智身科技展台,五大区域构成一条清晰的"干活能力"展示链路。 表演舞台区三大亮点集中呈现:智身银毅NE01人形机器人上演《功夫女足》多机协同舞蹈,单关节峰值扭矩达 180N・m;智身钢镚 L2 四足机器人展示镂空楼梯感控一体通行与 7:1 极限负载自重比。 台面操作演示区,智身科技带来首款行业级可组装教育四足机器人,从通识启蒙、专业实训到工程应用,支撑不同阶段的具身智能教育建设;MATRiX 2.0开放世界仿真训练平台提供手柄虚拟对战体验,打通从虚拟训练到真实部署的链路,现已开源。 地面操作演示区,智身铜锤M1搭载铁臂TA01机械臂展示VLA操作能力,即使智身铜锤M1腿部、本体持续运动时,智身铁臂 TA01 末端依然能保持稳定。系统能够根据动作与负载变化实时协…
Insider Brief Generalist AI has released GEN-1.5, a new robot foundation model designed to learn physical tasks from one or a few demonstrations while also generalizing to some new situations without task-specific training. The company said GEN-1.5 can learn some tasks after being shown a single three- to 12-second demonstration, without updating the model’s underlying […]
8月19日,以"人机共生、产需共融"为主题的2026世界机器人大会在北京正式开幕。作为全栈具身智能大模型公司,超维动力携全球首个人形机器人自主乒乓球完整对局成果、SMASH 2.0高动态人形乒乓系统、KAI世界模型、全球最高117个全身自由度的KAIBot高拟人本体、KAI Hand超高自由度灵巧手与KAI Halo第一视角数采头环等全栈矩阵登场,为观众带来一场"看得见、玩得到"的具身智能硬核体验。 SMASH 2.0:全球首个自主完整对局,现场开放人机对打 此前,超维动力已正式发布全球首个人形机器人自主乒乓球完整对局,标志着高速动态场景下"感知—决策—控制"全链路自主能力的里程碑式突破。本届大会,超维动力在展台现场搭建标准乒乓球交互体验区,全面展示基于自主研发的新一代SMASH 2.0系统——从视觉感知、轨迹预测、动作规划到全身控制的闭环串联能力。 乒乓球速度极快、旋转多变、落点难测,留给机器人的反应时间仅在毫秒级,真正难的不只是挥臂,而是肩、肘、腕、腰、腿多关节在极短时间内的协同配合。SMASH 2.0融合高速视觉感知、实时轨迹预测、全身运动控制与具身智能决策,在毫秒级窗口完成来球…
8月19日,2026世界机器人大会在北京亦庄开幕。作为全球机器人产业的重要风向标,本届大会汇集了人形机器人、工业机器人及大量具身智能解决方案。在星炽动力展台,搭载PULSE具身智能系统的机器人正在进行多场景任务演示,观众可以直观看到同一套模型如何在不同环境中理解指令并调整执行策略。 对于星炽动力而言,此次参展的三个关键词——PULSE、多场景与真机数据——相互关联,共同指向一个问题:当机器人离开展馆,进入社区服务、教育科研、商业服务等日常环境,支撑它完成任务并持续适应的能力从哪里来? 从一台机器人,看多场景能力如何复用 展会呈现的是能力切面,真实场景面对的则是不断变化的空间、对象、流程与用户需求。对机器人企业而言,多场景并非简单的场景数量叠加,更考验底层能力的迁移效率。展会现场,星炽动力展示了同一机器人平台在社区配送、科研教学、商业导览等场景的应用案例。其思路是以统一的具身智能底座支撑机器人理解任务、规划行动,并根据环境变化调整执行方式。 这套能力的核心是星炽动力自主研发的PULSE具身智能大模型。按照星炽动力的定义,PULSE是一套由用户意图驱动的世界动作模型,它同时理解物理环境与用…
当机器人开始理解世界,通用智能才进入物理空间。 作者丨吴思梦 陈嘉欣 编辑丨岑 峰 过去几年,机器人解决的问题是“如何更精准地执行动作”。 在工业生产线上,机器人可以完成焊接、搬运、装配等大量重复任务,但这些能力建立在一个前提之上:环境、流程和目标都已经被提前定义。机器人并不需要真正理解世界,只需要按照预设路径完成任务。 但当机器人开始走出工厂,进入物流、服务甚至家庭场景,问题正在发生变化。面对一个没有见过的物体、一个没有训练过的任务,机器人能不能理解环境、预测变化,并找到完成任务的方法? 这成为通用机器人走向下一阶段必须回答的问题。 大模型的发展为机器人提供了新的可能。语言模型证明,人工智能可以通过大规模数据学习复杂规律,并获得跨任务泛化能力。但机器人面对的并不是语言世界,而是真实的物理世界。 因此,具身智能领域正在探索一个更底层的问题:大模型究竟应该如何赋能机器人? 目前行业的主流方向是 VLA。它让机器人通过学习大量人类行为数据,将视觉、语言和动作统一起来。但 VLA 的核心仍然是模仿——模型学习人类过去如何行动。 问题在于,如果机器人面对一个从未见过的新任务,仅依靠已有动作经验…