Physical reasoning and planning world model
V-JEPA 2
A video-trained latent world model for understanding, predicting, and planning actions in the physical world.
01 / Overview
Self-supervised video representation learning, action-conditioned prediction, physical reasoning, robot planning.
A video-trained latent world model for understanding, predicting, and planning actions in the physical world.
Describe released benchmark and robot results precisely rather than claiming general physical intelligence.
Use V-JEPA 2 to explain the planner lane, not to rank image quality against Marble or Echo.
02 / Strengths
Where it stands out
- Represents the understand-predict-plan branch of world models clearly.
- Open artifacts and benchmarks make the research inspectable.
- Demonstrates how a model can evaluate possible actions without rendering a consumer fantasy world.
03 / Boundaries
What not to overclaim
- No consumer create surface or prompt-to-3D output.
- Research robot demonstrations do not prove safe general robot control.
- Short-horizon benchmark results should not be generalized to open-ended physical intelligence.
A latent prediction and planning model, not a consumer application that renders explorable worlds.
04 / Access & output
How the model reaches people
Official Meta explanation plus open code and model checkpoints for researchers and developers.
- Platforms
- GitHub · Model checkpoints · Research benchmarks
- Outputs
- Latent video representations · Predicted state representations · Planning signals
- Inputs & compatibility
- Video inputs · Research robot-planning pipelines · Local ML environment
05 / Practical path
How to evaluate it
- 01
Read Meta's official overview to understand the encoder-predictor design.
- 02
Use the repository and model cards for exact checkpoints and environment requirements.
- 03
Keep physical-understanding benchmarks separate from consumer world-generation comparisons.
Primary evidence