xLLM: Light-weight infrastructure for express LLM pre-training
- Type
- code
- Venue
- GitHub (ifm-ai)
- Year
- 2026
- Source
- github
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Apache-2.0 PyTorch framework for long-context LLM pretraining supporting dense, MoE, and MoVA architectures, with FSDP1/FSDP2, data/model/context parallelism, custom CUDA kernels, FlashAttention, online tokenization and packing, evaluation, HuggingFace export, and vLLM support.
Keywords
pretraining · infrastructure · long-context · moe · cuda · open-source
Topics
pretraining, infrastructure, long-context, moe, cuda
Research notes
- Discovery: Posted in #random-papers on 2026-09-28; repo README fetched directly from GitHub on 2026-09-29.
- Method: PyTorch training framework: FSDP 1/2 + data/model/context parallelism, custom CUDA kernels, FlashAttention, online tokenization/packing for long-context runs.
- Limitations: Repo created 2026-09-02 (very new at time of posting); benchmark numbers and maturity were not verified. Requires PyTorch >=2.11 and CUDA >=12.8.
- From ifm-ai, the same org behind the K2/K2-Horizon models also appearing in this channel.