Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- Type
- other
- Venue
- arXiv / Meta FAIR
Summary
I-JEPA: a ViT context encoder plus a narrow predictor, conditioned on positional mask tokens, predicts EMA target-encoder patch representations of several large target blocks from one context block. Multi-block masking (4 targets scale 0.15–0.2; context 0.85–1.0 minus overlap). No view augs. ViT-H/14 IN1K 300 ep: IN linear 79.3 / 1% 73.3; ViT-H/16@448: 81.1 / 77.3. Beats MAE linear-probe; competitive with DINO/iBOT; Clevr/Dist 72.4 vs DINO 53.4. ViT-H/14 on 16 A100 <72h / <1200 GPU-h. ICCV 2023.
Keywords
i-jepa · jepa · ssl · vit · mae · meta-fair · iccv · masking · representation-learning
Topics
self-supervised learning, JEPA, vision transformers
Research notes
- Primary: arxiv abs (cs.CV; also cs.AI, cs.LG, eess.IV). License: arXiv.org perpetual non-exclusive on HTML at check. Assran McGill/Mila/Meta; Duval/Misra/Bojanowski/Vincent/Rabbat/Ballas Meta AI (FAIR); LeCun NYU/Meta. Correspondence massran@meta.com. No official code URL on abs. FAIR repo https://github.com/facebookresearch/ijepa (3,487 stars at check; GitHub license NOASSERTION) not treated as abs-linked. HF paper page 7 upvotes; official-looking linked models facebook/ijepa_vith14_1k, facebook/ijepa_vitg16_22k etc. not copied into hf_* fields. Discord posted PDF. Trains on ImageNet rather than a new hosted corpus, so no datasets_local row. License field left blank per catalog convention.