← Back to explorer

DeepSeek-V3 Technical Report

Type
other
Venue
arXiv / DeepSeek

Summary

671B-parameter DeepSeekMoE (37B active: 1 shared + 8 of 256 routed experts, node-limited M=4) with Multi-head Latent Attention, auxiliary-loss-free load balancing, and 1-depth multi-token prediction. FP8 mixed precision, DualPipe pipeline parallelism, trained on 14.8T tokens with 2048 H800 GPUs. Full training 2.788M H800 GPU-hours (pretrain 2664K / context 119K / post-train 5K; ~$5.576M at $2/GPU-hour); no irrecoverable loss spikes. Context 4K→32K→128K via YaRN. Chat: MMLU 88.5, MMLU-Pro 75.9, GPQA-Diamond 59.1, MATH-500 90.2, AIME 2024 39.2, LiveCodeBench CoT 40.5, Arena-Hard 85.5. Code https://github.com/deepseek-ai/DeepSeek-V3.

Keywords

deepseek-v3 · moe · mla · mtp · fp8 · dualpipe · load-balancing · 671b

Topics

MoE, MLA, FP8 training, large language models

Research notes

  • Primary: arxiv abs (cs.CL; also cs.AI). License: arXiv.org perpetual non-exclusive on HTML at check. Authors listed as DeepSeek-AI (research@deepseek.com) plus a long alphabetical contributor list on the abs. Code https://github.com/deepseek-ai/DeepSeek-V3 (HF githubRepo; 104,255 stars at check; MIT). HF paper page 88 upvotes; official linked models deepseek-ai/DeepSeek-V3, DeepSeek-V3-0324, DeepSeek-V3-Base (plus many unofficial) not copied into hf_* fields. Unofficial linked datasets not substantial, so no datasets_local row. Discord posted HTML v1 §2. Pretraining mix is not a new public corpus. License field left blank per catalog convention.