← Back to explorer

xLLM: Light-weight infrastructure for express LLM pre-training

Type
code
Venue
GitHub (ifm-ai)
Year
2026
Source
github
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Apache-2.0 PyTorch framework for long-context LLM pretraining supporting dense, MoE, and MoVA architectures, with FSDP1/FSDP2, data/model/context parallelism, custom CUDA kernels, FlashAttention, online tokenization and packing, evaluation, HuggingFace export, and vLLM support.

Keywords

pretraining · infrastructure · long-context · moe · cuda · open-source

Topics

pretraining, infrastructure, long-context, moe, cuda

Research notes

  • Discovery: Posted in #random-papers on 2026-09-28; repo README fetched directly from GitHub on 2026-09-29.
  • Method: PyTorch training framework: FSDP 1/2 + data/model/context parallelism, custom CUDA kernels, FlashAttention, online tokenization/packing for long-context runs.
  • Limitations: Repo created 2026-09-02 (very new at time of posting); benchmark numbers and maturity were not verified. Requires PyTorch >=2.11 and CUDA >=12.8.
  • From ifm-ai, the same org behind the K2/K2-Horizon models also appearing in this channel.