WoW: Towards a World omniscient World model Through Embodied Interaction

Published in arXiv preprint (under review), 2025

WoW Model Architecture

Overview

Current video models like Sora rely on passive observation and struggle with physical causality. We hypothesize that authentic physical intuition must be grounded in extensive, causally rich interactions with the real world.

Key Contributions

  • Large-Scale Embodied Data: WoW is a 14B-parameter generative world model trained on 2 million real-world robot interaction trajectories.

  • SOPHIA Agent: A vision-language agent that evaluates DiT-generated outputs and iteratively refines language instructions toward physical realism.

  • Inverse Dynamics Model: Translates refined plans into executable robotic actions, closing the imagination-to-action loop.

  • WoWBench: A new benchmark for physical consistency and causal reasoning, where WoW achieves state-of-the-art in both human and autonomous evaluations.

  • Physical Intuition: Strong performance in physical causality, collision dynamics, and object permanence.

arXivProject Page

Recommended citation: Xiaowei Chi*, Peidong Jia*, Chun-Kai Fan*, ... Zhiyuan Jiang, ... et al. (2025). "WoW: Towards a World omniscient World model Through Embodied Interaction." arXiv preprint arXiv:2509.22642.
Download Paper