Embodied Intelligence Observer

Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

Research

Source: arXiv cs.ROPublish time unverified

arXiv:2608.26103v1 Announce Type: new Abstract: Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update. This form of in-context learning (ICL) turns generalization into a problem of task specification.

Multi-source coverage2

View full event →
Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization | Embodied Intelligence Observer