← Back to explorer

Cramming: Training a Language Model on a Single GPU in One Day

Type
other
Venue
arXiv / University of Maryland

Summary

Cramming: train a transformer MLM from scratch on one consumer GPU in 24h with no pretrained checkpoints. Scaling laws still hold at this budget, so architecture swaps barely move loss; gains come from faster steps at similar size (PreNorm, no QKV/FFN biases, sparse MLM head), a one-cycle LR, dropout-off, and filtered/sorted C4. A6000 1-day GLUE-dev 78.6 vs fully trained BERT-base 80.9 (MNLI 83.9/84.1 vs 83.2/83.4). Code https://github.com/JonasGeiping/cramming.

Keywords

cramming · bert · mlm · scaling-laws · glue · single-gpu · umd

Topics

efficient pretraining, BERT, scaling laws

Research notes

  • Primary: arxiv abs (cs.CL; also cs.LG). License: arXiv.org perpetual non-exclusive on HTML at check. Comment: 22 pages; code at https://github.com/JonasGeiping/cramming (HF githubRepo; 1,368 stars at check; no LICENSE file at check). UMD; correspondence jgeiping@umd.edu, tomg@umd.edu. HF paper page 0 upvotes; official-looking linked models JonasGeiping/crammed-bert and crammed-bert-legacy not copied into hf_* fields. Linked HF tokenized-Pile shards (the_pile_WordPiecex32768_*) are processed public Pile, not a new hosted corpus, so no datasets_local row. Discord posted PDF. License field left blank per catalog convention.