H-JEPA · Diagram study

Inverse dynamics in H-JEPA

Each level gets an inverse-dynamics head that recovers the action from two consecutive states, keeping the controllable robot in the latent on real robot video.

Inverse-dynamics heads at each level of H-JEPA Left: the two-level hierarchy built from the trajectory o0, a0, o1, a1, o2, a2, o3, a3, o4, with purple arcs marking each pair of consecutive states and the action between them. At level 1 every consecutive pair is used, including states that do not reach level 2. At level 2 the pairs are z0, z1 with macro-action a0, and z1, z2 with macro-action a1. Right: at each level, an inverse-dynamics head I takes two consecutive latent states and predicts the action; the prediction is regressed onto the stop-gradient action target. The resulting loss is added to the per-level objective with weight gamma. Two levels · stride s₂ = 2 · state window w₂ = 1 Build the hierarchy Inverse dynamics at each level LEVEL 2 LEVEL 1 o_{0} E^{(1)} z_{0}^{(1)} a_{0} A^{(1)} a_{0}^{(1)} o_{1} E^{(1)} z_{1}^{(1)} a_{1} A^{(1)} a_{1}^{(1)} o_{2} E^{(1)} z_{2}^{(1)} a_{2} A^{(1)} a_{2}^{(1)} o_{3} E^{(1)} z_{3}^{(1)} a_{3} A^{(1)} a_{3}^{(1)} o_{4} E^{(1)} z_{4}^{(1)} Observation–action trajectory E^{(2)} z_{0}^{(2)} E^{(2)} z_{1}^{(2)} E^{(2)} z_{2}^{(2)} A^{(2)} a_{0}^{(2)} A^{(2)} a_{1}^{(2)} For each level ℓ = 1, 2 z_{t}^{(ℓ)} z_{t+1}^{(ℓ)} I^{(ℓ)} â_{t}^{(ℓ)} ℒ_{idm}^{(ℓ)} sg[a_{t}^{(ℓ)}] Consecutive states Predicted action Target action stop-gradient Regress the action from two states ‖I^{(ℓ)}(z_{t}^{(ℓ)}, z_{t+1}^{(ℓ)}) − sg[a_{t}^{(ℓ)}]‖_{2}^{2} averaged over t and action dims Per-level objective ℒ^{(ℓ)} = ℒ_{pred}^{(ℓ)} + λ_{ℓ} SIGReg(Z^{(ℓ)}) + γ_{ℓ} ℒ_{idm}^{(ℓ)} At level 2, the target is the pooled macro-action from A⁽²⁾.

Encode each observation and primitive action at level 1, then form level-2 states and pooled macro-actions.

0:00 / 0:14
Latent statesEncoded actionsTraining objectives and IDM

E denotes a state encoder, A an action encoder, F a predictor, and I an inverse-dynamics head. sg is the stop-gradient on the action target, so the IDM gradient reaches the encoder only through the two states. Without this term, a purely predictive objective on DROID can keep the static background and drop the robot.