OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining
Abstract
OpenWAM factorizes world-action pretraining into modular components to identify key design principles, yielding a scalable open model with strong simulation and real-robot performance.
World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tightly coupled, obscuring which design choices matter and why. We introduce OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program. OpenWAM-Infra factorizes the WAM design space into composable modules with unified training, inference, deployment, and evaluation. On this substrate, OpenWAM-Study examines three questions through controlled experiments: what to inherit, how world and action learning interact, and how their synergy scales; and distills three principles: upstream knowledge transfers through a sufficiently capable generative backbone and a compact, information-rich latent space; world-action synergy requires dedicated action capacity, explicit world-to-action information flow, and synchronized joint denoising; and embodied pretraining principally improves out-of-domain generalization, with one-stage co-training over egocentric and robot data integrating world coverage and action grounding. Composing these principles, we build OpenWAM-ฮฑ, an open WAM pretrained on roughly 6,400 hours of egocentric human and robot data and evaluated across simulation and real-world benchmarks. Across the eight simulation benchmarks and the real-robot experiments, which together span embodiments from single-arm and bimanual manipulation to dexterous hands, OpenWAM-ฮฑ delivers consistently excellent performance, sustaining its top-tier standing from simulation to the physical world. We release the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.
Community
We introduce OpenWAM, an open, modular exploration towards systematic world-action model pretraining. OpenWAM includes three parts:
๐งฑ OpenWAM-Infra: We factorizes the WAM design space into composable modules, with unified training, inference, deployment, and evaluation across 8 simulation benchmarks.
๐ฌ OpenWAM-Study: Through controlled experiments, we try to answer three questions: what upstream knowledge to inherit, how to create world and action learning synergy, and how their synergy scales.
๐ OpenWAM-ฮฑ: Building on the insights from OpenWAM-Study, we pretrain a fully-open WAM on ~6,400 hours (518.5M frames) of egocentric human and robot data, yielding excellent performance across various simulation and real-world tasks.
๐ Results
- ๐ Top tier across 8 simulation benchmarks spanning 5 embodiments, with SOTA on the mobile bimanual benchmark EBench
- ๐ #1 on the real-world bimanual RoboDojo-Real leaderboard, doubling pi0.5's success rate
- ๐ฅ Consistently outperforms representative WAM and VLA baselines on single-arm and dexterous-hand real robots
We hope OpenWAM-ฮฑ serves as is a strong, reproducible baseline for future work.
๐ฆ Fully open: infra code and pretrained/post-trained weights.
๐ Project page: https://openwam-official.github.io/
๐ป Code: https://github.com/OpenWAM-Official/OpenWAM
๐ค Models & data: https://huggingface.co/OpenWAM
Models citing this paper 46
OpenWAM/OpenWAM-Alpha-Real-Dexterous-Hand-Wuji
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper