>10x More Efficient Pretraining
- Type
- blog
- Venue
- Magic blog
- Year
- 2026
- Source
- web
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Magic's research update on compute-efficient pretraining toward trillion-parameter models. Their pretraining recipe — the multiplicative result of tens of changes across model architecture, optimizer, training objective, and data, plus fixing minor bugs — matches DeepSeek V4 Pro Base quality using ~50x fewer FLOPs (roughly half of GPT-3's pretraining compute, ~$0.5M on GB200). Scaling the same recipe 10x (~$4M) meaningfully outperformed all publicly available open-weight base models on perplexity evaluations; by their fitted scaling laws, training an equally capable model under DeepSeek V4 Pro's recipe would cost >$100M. Magic frames pretraining plus agentic RL plus long context as sufficient for superhuman coding agents; next steps are scaling RL with long context, RL against the model's own latent knowledge of its intent, and further pretraining improvements.
Keywords
pretraining · scaling laws · efficiency · Magic · compute
Topics
pretraining, scaling laws, efficiency, Magic, compute
Research notes
- Method: Compounded algorithmic-efficiency recipe across architecture, optimizer, objective, and data; bits-per-byte loss on heldout data (private code repos, heldout research papers, heldout math reasoning) used to fit scaling laws projecting compute needed for a given capability; a short math RL run from base as a post-RL sanity check.
- Key findings: DeepSeek V4 Pro Base quality matched at ~50x fewer FLOPs (~$0.5M on GB200); 10x scale-up (~$4M) beat all public open-weight base models on perplexity evals
- Limitations: No model weights released; no external researcher has independently run the recipes. Evaluation methodology designed and executed entirely in-house (Fireworks provided only a partial cross-check on baseline log probabilities). Whether the efficiency gains are scale-invariant or scale-dependent is unresolved — MIT FutureTech research (Gundlach et al., ICML 2026) found most pretraining efficiency innovations are scale-dependent, which would make the gains largest at frontier scales rather than democratizing.
- Blog page not fetched directly (rate-limited); substance verified via the magic.dev page snippet in search results, Magic's LinkedIn post, and techtimes coverage. Exact publication date not verified.