← Back to explorer

NITP

Type
repo
Venue
aHapBean (GitHub)
Year
2026
Source
github
Access
free
Added
2026-08-14T18:45:00Z
Verified
2026-08-14T18:45:00Z

Summary

Landing page for Next Implicit Token Prediction (NITP, ICML 2026): keep NTP and add a cosine-alignment loss so the projected final hidden state at t predicts a stop-gradient shallow-layer representation of token t+1 (the implicit token), with the projection head discarded at inference. README: denser latent supervision without extra data, encoders, or backbone passes; higher effective rank and less last-state cosine collapse vs NTP. MoE 1.9B-A0.3B to 9B-A1B (9B-A1B average 40.27 to 42.94; MMLU-Pro / C3 / CommonsenseQA / ARC-Challenge / GSM8k gains) and dense 0.5B-3B; frozen 3B MoE mean-pooled last-layer reps improve 23/25 English MTEB tasks (overall 39.24 to 41.56). About 2.3% training FLOPs / 1.8% wall-clock in a 5k-step 9B MoE run; zero inference overhead. Implementation code still marked coming soon.

Keywords

nitp · next-token-prediction · pretraining · implicit-token · icml-2026 · moe · representation-geometry

Topics

LLM pre-training, representation learning

Research notes

  • Primary: GitHub README + API (license null / no LICENSE file; language null; 35 stars / 3 forks at check). Paper arXiv 2605.24956 (cs.CL; HTML arXiv perpetual non-exclusive license; comment Accepted at ICML 2026). Contact zhangxiangdong@sjtu.edu.cn. Discord posted the repo. README says implementation code coming soon (repo is assets + paper landing page). Not a new hosted corpus, so no datasets_local row.