FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution
- Type
- paper
- Venue
- arXiv (cs.DC)
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
An edge-native mixture-of-experts serving system that treats a personal machine as a unified, elastic inference platform rather than a small GPU. FreeToken co-designs model layout and loading, expert residency, CPU-GPU execution, agentic state reuse, and runtime memory management around the realities of local AI: agentic workloads with shifting execution patterns and heterogeneous per-machine hardware. Instead of a fixed offloading strategy, it continuously maps computation and model state onto available resources.
Keywords
inference · MoE · edge · serving · systems · open weights
Topics
inference, MoE, edge, serving, systems
Research notes
- Method: Bandwidth-adaptive execution: full serving-stack co-design (model layout/loading, expert residency, CPU-GPU execution, agentic state reuse, runtime memory management) with continuous remapping of compute and state to observed hardware resources.
- Key findings: Supports 20+ MoE models and real coding/tool-using agents from an 8GB laptop GPU to a single workstation GPU; Changes what local machines can serve: 35B model on a laptop, 284B on a gaming desktop, 753B GLM-5.2 on a single workstation GPU
- Limitations: Evaluation is on the authors' hardware sweep; performance on other edge configurations and non-MoE architectures not characterized in the abstract.
- Authors state the system is released at flashml.ai.