AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Asks whether a world model can drive without training a driving policy. Systematically evaluates existing action-conditioned JEPA world models (LeWM, DINO-WM, JEPA-WM) for end-to-end autonomous driving in a goal-conditioned zero-shot planning setting that uses ground-truth future observations as goals, isolating world-model quality from policy learning. Finds existing JEPA world models are either accurate for driving but computationally expensive, or efficient but insufficient for planning. Proposes AD-E2E-JEPA with a SIGReg-regularized learnable projector on projected patch embeddings, cutting planning patches 16x and embedding dimension 4x for a 100x inference speedup at retained planning performance (0.8 s for an 8-frame rollout over 256 candidate trajectories). Without any driving policy, the world model alone reaches goals 20 m away within 4.0/2.8 m displacement using 256/8,192 candidate trajectories. On NAVSIMv2: 67.3/72.9 EPDMS with multiplicative safety metrics, 84.1/86.5 without them. The self-supervised pretrained projector also lifts downstream imitation learning from 80.2 to 85.4 EPDMS.
Keywords
autonomous driving · world models · JEPA · zero-shot planning · NAVSIMv2
Topics
autonomous driving, world models, JEPA, end-to-end driving
Research notes
- Discovery: announced by Haoran Zhu (@2512185195Zhu) in an X thread on 2026-09-29 (first post of a 7-part thread): https://x.com/2512185195Zhu/status/2104961832393855073
- Code: https://github.com/HaoranZhuExplorer/AD-E2E-JEPA
- Joint work with Kevin Zhang (@kevinghstz), Yann LeCun (@ylecun), and Anna Choromanska. Submitted to arXiv 2026-09-28.
- Position: JEPA world-model approach to end-to-end autonomous driving without learned driving policies.