← Back to explorer

PithTrain: A Compact and Agent-Native MoE Training System

Type
paper
Venue
arXiv / Carnegie Mellon University / NVIDIA
Year
2026
Source
arxiv
Access
free
Language
en
Added
2026-08-14T16:58:34Z
Verified
2026-08-14T16:58:34Z

Summary

Defines agent-task efficiency (ATE) and ships PithTrain, an ~11K-line Python-native MoE stack (PP, FSDP DP, CP, EP, DualPipeV overlap, torch.compile, FP8) built on compactness, no implicit indirection, and in-repo agent skills. Matches or exceeds Megatron-LM tokens/s on GPT-OSS-20B, Qwen3-30B-A3B, and DeepSeek-V2-Lite on H100/B200 (e.g. 280.0K vs 264.1K tok/s on 4×8 H100 Qwen3-30B-A3B PP4/EP8). ATE-Bench (12 Q&A, 4 operate/profile, 4 new-feature ports) with Claude Opus 4.7: up to 62% fewer agent turns and 64% less active GPU time vs production frameworks on the hardest new-feature tasks. Code Apache-2.0 https://github.com/mlc-ai/pith-train.

Keywords

pithtrain · moe · agent-native · megatron · ate-bench · cmu · mlc

Topics

MoE training systems, coding agents

Research notes

  • Primary: arxiv abs (CC BY 4.0, cs.LG/AI/CL/DC). Code Apache-2.0 https://github.com/mlc-ai/pith-train (332 stars at check). CMU with NVIDIA (Shao, Chen). HF paper page 1 upvote; no linked models/datasets. ATE-Bench is a task suite over existing frameworks, not a standalone public corpus, so no datasets_local row. Discord posted abs link.