Researchers at Yann LeCun’s Advanced Machine Intelligence (AMI) Labs and academic institutions have developed H-JEPA, a world-model architecture that learns to plan at different timescales using separate representations of the environment.

Compared to flat JEPA world models, H-JEPA improves accuracy on navigation tasks while requiring less computation for planning.

The work addresses a challenge in developing AI systems that can understand and act in the physical world. For example, a robot navigating a building must decide where to go while also controlling its movements moment by moment. H-JEPA addresses this by planning across different levels of abstraction, allowing the system to reason about its destination at a high level and progressively translate its plans into precise physical actions.

Why JEPA world models need hierarchy

Joint-Embedding Predictive Architectures (JEPAs) learn to model the world without predicting every low-level detail. In contrast to generative models, which try to predict granular details such as pixels, JEPA models encode observations such as images or videos into latent representations, or compact sets of features that capture information about the scene. This makes them much more efficient than generative models and highly accurate at tasks such as navigating physical environments or manipulating objects, which don’t require generating pixels.

V-JEPA

V-JEPA architecture (source: Meta)

In action-conditioned JEPAs, the model receives actions in addition to visual data. It can simulate the consequences of different action sequences and compare the predicted outcome with a goal. 

Existing approaches face two related problems. Predicting a distant outcome may require simulating many small steps, which increases compute costs and can cause small errors to cascade. At the same time, the model must track fine-grained physical dynamics while also evaluating its progress toward a long-term objective in the same representation.

Consider a four-legged robot navigating a maze. Its leg positions matter when controlling its movements. But a planner choosing a route primarily needs the robot’s location. Encoding both the robot's location and granular physical movements in the same latent representation makes tasks complicated and difficult to learn.

H-JEPA solves this problem by learning representations at different levels. Long-range planning can work with different information from the model that handles immediate actions.

How H-JEPA learns and plans

H-JEPA is trained on sequences of observations and the actions that connect them. The lowest level encodes visual observations and predicts the next latent state given preceding states and actions. Higher levels take representations from the level below and predict transitions across longer time intervals.

Each level has its own state encoder, action encoder, and predictor. The state encoder represents the environment at that timescale. The action encoder summarizes actions over the corresponding interval. The predictor learns how the represented state changes.

H-JEPA architecture

H-JEPA architecture (source: H-JEPA project page)

The researchers trained the hierarchy end to end. Each level compares its predicted representation with the observed future representation. A regularization mechanism prevents the model from producing the same representation for every observation. Training signals from higher levels also flow through lower-level encoders.

The intervals between levels are predefined, but the information their representations retain is learned. When parts of an environment change at different speeds, higher levels can preserve slower features that remain predictable over longer periods while discarding fast-changing details.

In the maze example, the lowest level needs the robot’s joint configurations to predict its immediate movements. A higher level can preserve the robot’s location without retaining all that detail.

During planning, H-JEPA encodes the current observation and target observation at every level. The high-level planner searches for actions whose predicted outcomes are closer to the goal in its latent space.

Its predicted states become subgoals for the next level, which searches for actions to reach them. Each level refines the plan until the lowest level produces primitive actions for execution.

H-JEPA planning process

H-JEPA planning process (source: H-JEPA project page)

In the maze, the top level can focus on progressing toward the destination while lower levels handle shorter transitions and physical movements. 

After executing a short sequence of actions, the system updates its plan based on new observations and the deviations it spots from its predicted trajectory.

H-JEPA in action

The researchers evaluated H-JEPA on four simulated navigation and manipulation environments: FourRoom Distractors, Visual AntMaze, Push-T, and OGBench Cube. They compared it with LeWM, a single-level JEPA, and HWM, another world model that plans hierarchically within a shared latent space.

In FourRoom, AntMaze and Cube, adding levels up to a three-level hierarchy generally improved success while reducing planning compute. The advantage of learning separate representations was strongest in AntMaze, where three-level H-JEPA achieved nearly twice HWM’s success rate.

H-JEPA results on industry benchmarks

H-JEPA results on industry benchmarks (source: H-JEPA project page)

In AntMaze, the robot’s position remains recoverable from higher-level representations, while information about its leg configuration fades. Higher levels also provide smoother estimates of distance to a goal along maze corridors, giving the planner a clearer signal for navigating around walls.

The team then tested H-JEPA on DROID, a dataset of real robot manipulation videos with varying objects, lighting, and backgrounds. Pure predictive training tended to retain static scene details and lose information about the moving robot. They then added an “inverse dynamics” objective, which predicts the action connecting two observed states, and it helped preserve action-relevant information. H-JEPA improved offline planning fidelity over the baselines with less planner compute. 

H-JEPA could be relevant to robotic navigation, warehouse automation, and manipulation systems that must coordinate long-term goals with precise physical movements. However, the model is not without its limits. Experiments show that more hierarchy is not always better, and the benefits depend on the scope of the training data. 

The researchers propose connecting higher-level representations to language, which could eventually let agents receive goals as instructions rather than target images. For now, H-JEPA shows the potential of learning different representations for immediate movements and distant objectives.