H-JEPA · Animation study

Learning from an interleaved trajectory

An alternate view with observations and actions in temporal order, from o₀ through o₄.

Hierarchical training with interleaved observations and actions The input trajectory alternates o0, a0, o1, a1, o2, a2, o3, a3, o4. Level 1 encodes all five observations and four actions. With stride two and state window one, level 2 encodes states at times zero, two, and four into three states. Its first macro-action aggregates action embeddings zero and one; its second aggregates two and three. Every level trains with predictive loss and SIGReg, with gradients flowing through the hierarchy. Two levels · stride s₂ = 2 · state window w₂ = 1 Build the hierarchy Train each level LEVEL 2 LEVEL 1 o_{0} E^{(1)} z_{0}^{(1)} a_{0} A^{(1)} a_{0}^{(1)} o_{1} E^{(1)} z_{1}^{(1)} a_{1} A^{(1)} a_{1}^{(1)} o_{2} E^{(1)} z_{2}^{(1)} a_{2} A^{(1)} a_{2}^{(1)} o_{3} E^{(1)} z_{3}^{(1)} a_{3} A^{(1)} a_{3}^{(1)} o_{4} E^{(1)} z_{4}^{(1)} Observation–action trajectory E^{(2)} z_{0}^{(2)} E^{(2)} z_{1}^{(2)} E^{(2)} z_{2}^{(2)} A^{(2)} a_{0}^{(2)} A^{(2)} a_{1}^{(2)} For each level ℓ = 1, 2 z_{t}^{(ℓ)} F^{(ℓ)} ẑ_{t+1}^{(ℓ)} ℒ_{pred}^{(ℓ)} a_{t}^{(ℓ)} z_{t+1}^{(ℓ)} Context Prediction Action Target Regularize encoded states SIGReg(Z^{(ℓ)}) Z: encoded states across the batch and time Per-level objective ℒ^{(ℓ)} = ℒ_{pred}^{(ℓ)} + λ_{ℓ} SIGReg(Z^{(ℓ)}) Sum the objectives and train all levels jointly.

Encode each observation and primitive action into its own latent representation.

0:00 / 0:14
Latent statesEncoded actionsTraining objectives

Adapted from the original training TikZ. E denotes a state encoder, A an action encoder, and F a predictor. The same construction extends to deeper hierarchies and larger state windows.