H-JEPA · Diagram study
Inverse dynamics in H-JEPA
Each level gets an inverse-dynamics head that recovers the action from two consecutive states, keeping the controllable robot in the latent on real robot video.
Encode each observation and primitive action at level 1, then form level-2 states and pooled macro-actions.
0:00 / 0:14
Latent statesEncoded actionsTraining objectives and IDM
E denotes a state encoder, A an action encoder, F a predictor, and I an inverse-dynamics head. sg is the stop-gradient on the action target, so the IDM gradient reaches the encoder only through the two states. Without this term, a purely predictive objective on DROID can keep the static background and drop the robot.