← Back to explorer

OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration

Type
other
Venue
arXiv / SJTU / Alibaba Qwen / UW–Madison / UIUC / Mila

Summary

OPUS (Optimizer-induced Projected Utility Selection) linearizes AdamW/Muon one-step updates, sketches ghost outer-product gradients with CountSketch, and Boltzmann-samples a batch aligned with a Bench-Proxy retrieved from the corpus. Scoring overhead 4.7% vs naive 3.5×. GPT-2 XL Muon FineWeb 30B: avg 41.75 vs random 40.29, beating 60B random (41.29) and static filters. FineWeb-Edu: OPUS on score-3 beats baselines trained on score 4+5 (XL 44.99). Qwen3-8B-Base CPT on SciencePedia: 0.5B tokens beats 3B full CPT. Code https://github.com/gszfwsb/OPUS.

Keywords

opus · data-selection · muon · adamw · fineweb · qwen3 · sciencepedia · icml · ghost-gradients · countsketch · sjtu · alibaba-qwen

Topics

data selection, LLM pretraining, optimizers

Research notes

  • Primary: arxiv abs (cs.CL). CC BY 4.0 on HTML. EPIC Lab SJTU / Qwen Team Alibaba / UW–Madison / UIUC / Mila. Wang intern at Qwen; Wang/Ouyang/Xu/Hu/Liu equal contrib. Correspondence xingzhang.rxz / liudayiheng.ldyh@alibaba-inc.com, zhanglinfeng@sjtu.edu.cn. Code MIT https://github.com/gszfwsb/OPUS (29 stars / 4 forks at check); GitHub README frames as ICML 2026 Oral. HF paper page 355 upvotes, org Qwen; no linked models/datasets. Discord posted PDF. Uses public FineWeb/FineWeb-Edu/SciencePedia; no new corpus, so no datasets_local row. License field left blank per catalog convention.