PithTrain: A Compact and Agent-Native MoE Training System
- Type
- paper
- Venue
- arXiv / Carnegie Mellon University / NVIDIA
- Year
- 2026
- Source
- arxiv
- Access
- free
- Language
- en
- Added
- 2026-08-14T16:58:34Z
- Verified
- 2026-08-14T16:58:34Z
Summary
Defines agent-task efficiency (ATE) and ships PithTrain, an ~11K-line Python-native MoE stack (PP, FSDP DP, CP, EP, DualPipeV overlap, torch.compile, FP8) built on compactness, no implicit indirection, and in-repo agent skills. Matches or exceeds Megatron-LM tokens/s on GPT-OSS-20B, Qwen3-30B-A3B, and DeepSeek-V2-Lite on H100/B200 (e.g. 280.0K vs 264.1K tok/s on 4×8 H100 Qwen3-30B-A3B PP4/EP8). ATE-Bench (12 Q&A, 4 operate/profile, 4 new-feature ports) with Claude Opus 4.7: up to 62% fewer agent turns and 64% less active GPU time vs production frameworks on the hardest new-feature tasks. Code Apache-2.0 https://github.com/mlc-ai/pith-train.
Keywords
pithtrain · moe · agent-native · megatron · ate-bench · cmu · mlc
Topics
MoE training systems, coding agents
Research notes
- Primary: arxiv abs (CC BY 4.0, cs.LG/AI/CL/DC). Code Apache-2.0 https://github.com/mlc-ai/pith-train (332 stars at check). CMU with NVIDIA (Shao, Chen). HF paper page 1 upvote; no linked models/datasets. ATE-Bench is a task suite over existing frameworks, not a standalone public corpus, so no datasets_local row. Discord posted abs link.