NITP
- Type
- repo
- Venue
- aHapBean (GitHub)
- Year
- 2026
- Source
- github
- Access
- free
- Added
- 2026-08-14T18:45:00Z
- Verified
- 2026-08-14T18:45:00Z
Summary
Landing page for Next Implicit Token Prediction (NITP, ICML 2026): keep NTP and add a cosine-alignment loss so the projected final hidden state at t predicts a stop-gradient shallow-layer representation of token t+1 (the implicit token), with the projection head discarded at inference. README: denser latent supervision without extra data, encoders, or backbone passes; higher effective rank and less last-state cosine collapse vs NTP. MoE 1.9B-A0.3B to 9B-A1B (9B-A1B average 40.27 to 42.94; MMLU-Pro / C3 / CommonsenseQA / ARC-Challenge / GSM8k gains) and dense 0.5B-3B; frozen 3B MoE mean-pooled last-layer reps improve 23/25 English MTEB tasks (overall 39.24 to 41.56). About 2.3% training FLOPs / 1.8% wall-clock in a 5k-step 9B MoE run; zero inference overhead. Implementation code still marked coming soon.
Keywords
nitp · next-token-prediction · pretraining · implicit-token · icml-2026 · moe · representation-geometry
Topics
LLM pre-training, representation learning
Research notes
- Primary: GitHub README + API (license null / no LICENSE file; language null; 35 stars / 3 forks at check). Paper arXiv 2605.24956 (cs.CL; HTML arXiv perpetual non-exclusive license; comment Accepted at ICML 2026). Contact zhangxiangdong@sjtu.edu.cn. Discord posted the repo. README says implementation code coming soon (repo is assets + paper landing page). Not a new hosted corpus, so no datasets_local row.