CometVLA: Co-Training on an Embodied Data Pyramid towards Physical Understanding
Research
Source: arXiv cs.ROPublish time unverified
arXiv:2608.30289v1 Announce Type: new Abstract: Vision-language-action (VLA) models remain brittle in manipulation tasks that require physical commonsense. Current physical VQA data is typically disembodied and misaligned with robot action domains. Egocentric videos are used only as auxiliary pre-training. It remains unclear whether improved VLM physical understanding actually benefits downstream action generation. Therefore, we present CometVLA to close this gap.