Embodied Intelligence Observer

Latent Chain-of-Thought World Modeling for End-to-End Driving

Research

Source: arXiv cs.ROPublish time unverified

arXiv:2512.10226v3 Announce Type: replace-cross Abstract: Recent Vision-Language-Action (VLA) models for autonomous driving explore inference-time reasoning as a way to improve driving performance and safety in challenging scenarios. Most prior work uses natural language to express chain-of-thought (CoT) reasoning before producing driving actions. However, text may not be the most efficient representation for reasoning.

Latent Chain-of-Thought World Modeling for End-to-End Driving | Embodied Intelligence Observer