Hasty Briefsbeta

双语

OpenWAM: An Open Framework for Composable World-Action Models

5 hours ago
  • OpenWAM is a framework for composable world-action models that unifies prediction and control in robot learning.
  • It supports multiple interaction programs: video-then-action, action-then-video, joint denoising, and decoupled generation.
  • A 5B video expert (adapted from Wan2.2) is combined with a 2B action expert using a Mixture-of-Transformers (MoT) architecture.
  • Pretraining on 3.34M trajectories of robot and human interaction video provides a strong visual foundation.
  • Local-context inverse dynamics (IDM) and forward dynamics (FDM) models can be trained separately and composed at inference.
  • Counterfactual data (LIBERO-Long-CF) with 32,000 segments improves model robustness and transfer, especially for frozen IDM.
  • OpenWAM achieves high success on LIBERO benchmarks (e.g., 98.6% with VTA) and bimanual manipulation tasks (~92%).
  • Robot-video pretraining and MoT architecture significantly outperform naive baselines (e.g., 29.4-point gain on LIBERO-Long).
  • Frozen local-context IDM transfers to new tasks with 84.0% success when trained on counterfactual data (vs. 21.5% demo-only).
  • A unified checkpoint handles policy, IDM, and FDM but specialists remain more accurate; counterfactual data improves forward dynamics prediction.