Skip to content
Projects

Project

Spatio-Temporal World Model

Today's AI excels in digital domains but cannot understand the physical 4D world: robots see pixels, not space and time. We build a Spatio-Temporal World Model that fuses multi-modal sensory input into a unified 4D representation with physics consistency, enabling machines to perceive, remember, and predict dynamic environments.

Visualization for Spatio-Temporal World Model

Overview

Our central hypothesis: a world model can jointly represent 4D spatio-temporal states and predict physically consistent futures without pixel-level generation. The key result is a joint embedding architecture compressing multi-modal observations into structured world tokens, enabling latent-space prediction verified by a physics consistency module. This bridges perception and causal reasoning — allowing machines to move from observing the present to anticipating and explaining physical events.

Related Publications

Publication records will be added here as project outputs are released.

Impact Holders

Impact holders and user communities will be added as the project scope becomes clearer.