← Back to explorer

>10x More Efficient Pretraining

Type
blog
Venue
Magic blog
Year
2026
Source
web
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Magic's research update on compute-efficient pretraining toward trillion-parameter models. Their pretraining recipe — the multiplicative result of tens of changes across model architecture, optimizer, training objective, and data, plus fixing minor bugs — matches DeepSeek V4 Pro Base quality using ~50x fewer FLOPs (roughly half of GPT-3's pretraining compute, ~$0.5M on GB200). Scaling the same recipe 10x (~$4M) meaningfully outperformed all publicly available open-weight base models on perplexity evaluations; by their fitted scaling laws, training an equally capable model under DeepSeek V4 Pro's recipe would cost >$100M. Magic frames pretraining plus agentic RL plus long context as sufficient for superhuman coding agents; next steps are scaling RL with long context, RL against the model's own latent knowledge of its intent, and further pretraining improvements.

Keywords

pretraining · scaling laws · efficiency · Magic · compute

Topics

pretraining, scaling laws, efficiency, Magic, compute

Research notes

  • Method: Compounded algorithmic-efficiency recipe across architecture, optimizer, objective, and data; bits-per-byte loss on heldout data (private code repos, heldout research papers, heldout math reasoning) used to fit scaling laws projecting compute needed for a given capability; a short math RL run from base as a post-RL sanity check.
  • Key findings: DeepSeek V4 Pro Base quality matched at ~50x fewer FLOPs (~$0.5M on GB200); 10x scale-up (~$4M) beat all public open-weight base models on perplexity evals
  • Limitations: No model weights released; no external researcher has independently run the recipes. Evaluation methodology designed and executed entirely in-house (Fireworks provided only a partial cross-check on baseline log probabilities). Whether the efficiency gains are scale-invariant or scale-dependent is unresolved — MIT FutureTech research (Gundlach et al., ICML 2026) found most pretraining efficiency innovations are scale-dependent, which would make the gains largest at frontier scales rather than democratizing.
  • Blog page not fetched directly (rate-limited); substance verified via the magic.dev page snippet in search results, Magic's LinkedIn post, and techtimes coverage. Exact publication date not verified.