Yann LeCun's world-model startup, Advanced Machine Intelligence (AMI), has quietly shipped its first public research result. In a new paper called H-JEPA, researchers from AMI, New York University, INRIA Paris and Brown University describe a hierarchical world model built for visual planning - and, in keeping with the open stance LeCun set out when he left Meta, the team has released both the code and pretrained weights so anyone can load the model and reproduce the experiments.
The idea behind the paper is one LeCun has been articulating since 2022, when he was still Meta's chief AI scientist: complex goals need planning at multiple abstraction levels and time scales. H-JEPA stacks several JEPA (joint-embedding predictive architecture) layers, each of which predicts the state of the environment a different number of steps into the future, in its own representation space. When the system plans, the top layer sketches a rough route, hands the predicted intermediate states to the layer below as subgoals, and the process repeats downward until primitive, executable actions fall out.
The LeCun-flavored intuition is a trip from an NYU office to Paris: at the planning stage, what matters is the destination, the flight and the schedule - not the height of the curb outside the building. As you actually walk out the door, the relevant information flips: stairs, footing, which foot moves first. Robots face the same split. A quadruped that has reached roughly the right position may still look "far" from its goal frame because its legs are posed differently, which is exactly the kind of detail a single flat representation struggles to keep in its place.
On the paper's headline benchmark, the numbers are stark. In the Visual AntMaze navigation task, a three-level H-JEPA reached a 73.3 percent success rate, against 18.0 percent for the single-layer model the team compares against (LeWorldModel) - and the hierarchical version needed less planning compute to get there. Those are the researchers' own reported results on their own evaluation suite, not independently replicated, but they are large enough to matter.
They are also not uniform. In Push-T, a manipulation task that involves pushing a block, the three-level model did worse than the two-level one - more hierarchy is not automatically better. The team also ran an offline test on DROID, a dataset of real robot videos, where H-JEPA's predicted action trajectories tracked expert demonstrations more closely than the single-layer baseline. The evaluation suite spans FourRoom Distractors, Visual AntMaze, Push-T, OGBench Cube and DROID.
Architecturally, the paper's key move is separation. Earlier hierarchical attempts such as HWM made every level share one set of environment features; H-JEPA lets each level learn its own, and adds an inverse dynamics objective to keep action-relevant information from being squeezed out. In the maze experiments, the high levels retained position and corridor information while progressively discarding short-term details like leg posture - which is precisely the division of labor the theory calls for.
Why it matters: AMI is LeCun's bet that world models rather than ever-larger LLMs are the road to machines that plan and act, and H-JEPA is the first concrete, reproducible installment of that bet. A 73.3-versus-18.0 gap on one simulated maze does not settle anything about physical robots, and no outside lab has validated the numbers yet. But with code and weights public, the claim is now checkable - which is more than can be said for most closed frontier-lab demos.
Comments (0)
Log in to join the discussion
Log InNo comments yet