← Back to explorer

xLLM

Type
repo
Venue
GitHub
Year
2026
Source
github
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

PyTorch-based framework for long-context LLM pre-training. Supports dense, MoE, and MoVA transformer architectures; FSDP1/FSDP2, data/model/context parallelism; custom CUDA kernels, fused blocks, recomputation, and FlashAttention backends (with H200 benchmarks for K2 Horizon and Llama 3 8B). Includes an online data pipeline (parallel tokenization, async preparation, buffered shuffle, bestfit packing), checkpoint/resume, evaluation, and Hugging Face export + vLLM serving via xBridges.

Keywords

LLM training · training infra · long context · FSDP · context parallelism · PyTorch

Topics

LLM training infrastructure, long-context training

Research notes

  • Discovery: open-sourced by Xuezhe "Max" Ma (@MaxMa1987, Research Lead @ USC ISI / Asst. Prof. @ USC CS; PhD CMU) on 2026-09-28, quoting an @IFM_AI post on LLM pre-training/fine-tuning rework costs: https://x.com/MaxMa1987/status/2104650088610156753
  • Method: distributed training infra.
  • Apache 2.0.
  • Created 2026-09-02; 81 stars at research time.
  • Core contributors: @desai_Fan, @Chufan_Shi, @Shicheng_Wen, @_adharshbabu_, @BowenTan8, @IFM_AI, @USC_ISI.