← Back to explorer

Learn from your own latents and not from tokens: A sample-complexity theory

Type
other
Venue
arXiv / EPFL / University of Cambridge / Johns Hopkins University

Summary

On the Random Hierarchy Model (PCFG, depth L, m rules/symbol), supervised learning needs ~m^L samples and token-level SSL (MLM/diffusion) ~m^{L+1}. Latent prediction (cluster cousin context vectors level-by-level) recovers the non-root tree from ~v m^3 samples, independent of L up to logs. Confirmed with (i) iterative latent clustering, (ii) a stacked predictor-clusterer net trained by GD (also with local stop-grad rules), and (iii) the first sample-complexity analysis of data2vec, which implicitly does hierarchical latent prediction — so explicit H-JEPA stacking is largely redundant. Suggests own-latent SSL as a route to beat token-correlation scaling laws.

Keywords

jepa · data2vec · rhm · sample-complexity · latent-prediction · epfl · hierarchical-ssl · predictive-coding

Topics

self-supervised learning, JEPA, sample complexity, hierarchical data

Research notes

  • Primary: arxiv abs (cs.LG). CC BY 4.0 on abs HTML. EPFL / Cambridge DAMTP / JHU; equal contrib Korchinski/Favero. Correspondence daniel.korchinski@epfl.ch, af940@cam.ac.uk, matthieu.wyart@epfl.ch. No official code on abs. HF paper page 3 upvotes. Discord posted abs. Synthetic RHM only; no datasets_local row.