Researchers from Yann LeCun's AMI Labs, in collaboration with NYU, INRIA Paris, and Brown University, introduced H-JEPA, a hierarchical world model for long-horizon visual planning. The paper, published on arXiv on October 5, 2026, and the code were open-sourced, with weights available.
How does H-JEPA work?
Unlike flat JEPA models that plan in a single latent space, H-JEPA trains a hierarchy of action-conditioned JEPAs. Each higher level predicts further ahead in its own learned latent space. Planning proceeds top-down: the high level sets subgoals, and lower ones refine them into primitive actions. This discards fast, unpredictable details at higher levels, focusing on task-relevant states.
Results and impact
In simulations like Visual AntMaze, a three-level hierarchy raised the success rate from 18% (flat model) to 73%, with less planning compute. Tests in other navigation and manipulation environments confirmed gains. With inverse-dynamics supervision, it also improves offline planning on real DROID videos.
The code is on GitHub (kevinghst/H-JEPA), based on LeWM and stable-worldmodel. This advances LeCun's vision of world models as an alternative to generative LLMs. Sources: arXiv paper 2610.06805 and 36Kr report.
By GeekikiBot