OpenWAM: An Open Framework for Composable World-Action Models
5 hours ago
- OpenWAM is a framework for composable world-action models that unifies prediction and control in robot learning.
- It supports multiple interaction programs: video-then-action, action-then-video, joint denoising, and decoupled generation.
- A 5B video expert (adapted from Wan2.2) is combined with a 2B action expert using a Mixture-of-Transformers (MoT) architecture.
- Pretraining on 3.34M trajectories of robot and human interaction video provides a strong visual foundation.
- Local-context inverse dynamics (IDM) and forward dynamics (FDM) models can be trained separately and composed at inference.
- Counterfactual data (LIBERO-Long-CF) with 32,000 segments improves model robustness and transfer, especially for frozen IDM.
- OpenWAM achieves high success on LIBERO benchmarks (e.g., 98.6% with VTA) and bimanual manipulation tasks (~92%).
- Robot-video pretraining and MoT architecture significantly outperform naive baselines (e.g., 29.4-point gain on LIBERO-Long).
- Frozen local-context IDM transfers to new tasks with 84.0% success when trained on counterfactual data (vs. 21.5% demo-only).
- A unified checkpoint handles policy, IDM, and FDM but specialists remain more accurate; counterfactual data improves forward dynamics prediction.