01Known worlds
Games and metaverse taught the interface.
People know avatars, spaces, inventories, maps, shared places.
World model concept map
Use the world model concept map to connect AI video, spatial computing, digital twins, physical AI, and generated worlds.
Concept mapCore sentence
Minecraft and Roblox explain the mental model, the metaverse persistence and social space, Vision Pro spatial computing. EMO, Veo, Wan, Kling, and Ray explain controllable video; Cosmos and digital twins simulation.
Scene explainer
01Known worlds
People know avatars, spaces, inventories, maps, shared places.
02Generated media
Synthetic scenes became easy to watch and share, but behaved like clips.
03World models
Scenes remember space, respond to action, and support agents or simulation.
Concept flow
Past interface
Current surface
Spatial interface
Industrial layer
Core capability
Past interface
People already understood avatars, sandbox worlds, social rooms, user-built spaces.
Simple characters make presence easy: a person enters a world.
Minecraft and Roblox trained users to expect modifiable, shareable worlds.
Framed virtual worlds as social, persistent, identity-driven, despite manual tooling.
Current surface
The visible surface; deeper issues are control, consistency, and memory.
The same identity must move, emote, sing, and stay coherent.
Turn text, images, audio, and references into moving scenes.
Generated characters will need continuity across avatars and portraits.
Spatial interface
They are how generated worlds may be seen and operated.
Reframes computing as placed into space, not a flat screen.
Make real or imagined spaces computable.
Users feel located inside a generated or captured environment.
Industrial layer
Not entertainment; simulation for robots, vehicles, factories, and cities.
Model real systems so teams test changes before touching reality.
Models of how environments respond to motion, contact, and decisions.
Connect AI to real-world places, maps, and location-aware behavior.
Core capability
Modeling how worlds change under time, viewpoint, and action.
Preserve coherent state when the user moves, edits, or acts.
Reusable infrastructure for generating, predicting, and testing world states.
Agents need environments to observe, act, fail, and learn.
Bridge table
| Entry concept | Known for | Connects to | Meaning in world models |
|---|---|---|---|
| Blocky avatars / Minecraft | Simple identity inside a buildable world | Avatars, sandbox worlds, UGC | Generated worlds need persistent users, objects, and editable structure. |
| Metaverse | Persistent social virtual spaces | VR, Horizon Worlds, social identity | Automate world creation instead of relying on manual building. |
| Vision Pro | Spatial computing and immersive interface | AR, spatial video, 3D interaction | Generated worlds need spatial interfaces for viewing, editing, operation. |
| AI video | Generated motion, characters, and scenes | EMO, Veo 3.1, Wan2.7-Video, Kling, Ray | The video layer must become controllable, continuous, and stateful. |
| Digital twins | Simulation of real systems | Omniverse, robotics, LingBot-VA, LingBot-VLA, city and factory models | Useful when they predict and test real-world behavior. |
| World model | Predicting and generating world state | Genie 3, Marble, Cosmos, GWM-1 | Not a place or device; the model making worlds behave. |