Physical-Intelligence/openpi
Open-source code and fine-tuning pipeline for the π0/π0.5 VLA models from Physical Intelligence (OpenPI); a flagship VLA with streaming action outputs.
openvla/openvla
7B-parameter open-source VLA foundation model (Prismatic VLM trained on Open X-Embodiment data); fine-tunable on a single GPU and widely used as an embodied-model baseline.
NVIDIA/Isaac-GR00T
NVIDIA's open-source humanoid foundation model GR00T N1 (dual-system architecture: fast VLM system + slow diffusion action blocks), with data pipelines and fine-tuning scripts.
thu-ml/RoboticsDiffusionTransformer
Largest open-source bimanual manipulation diffusion foundation model (1B params, jointly trained across multiple robot embodiments); led by a Chinese team.
AgibotTech/Genie-Envisioner-V1
AgiBot's open embodied world-model and policy platform: a video-diffusion world model with instruction-conditioned streaming policies covering understanding, prediction and execution of manipulation (includes the GE-Sim world simulator).
Open-X-Humanoid/HEX
Whole-body humanoid VLA framework from the Tien Kung team (Humanoid-Aligned Experts): a vision-language-action model for cross-embodiment whole-body manipulation, with weights and data open on Hugging Face.
OpenGalaxea/GalaxeaVLA
Galaxea's open G0.5 vision-language-action model: a general manipulation foundation model with data and deployment tools.