← Back to explorer

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

Type
paper
Venue
arXiv (cs.DC)
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

An edge-native mixture-of-experts serving system that treats a personal machine as a unified, elastic inference platform rather than a small GPU. FreeToken co-designs model layout and loading, expert residency, CPU-GPU execution, agentic state reuse, and runtime memory management around the realities of local AI: agentic workloads with shifting execution patterns and heterogeneous per-machine hardware. Instead of a fixed offloading strategy, it continuously maps computation and model state onto available resources.

Keywords

inference · MoE · edge · serving · systems · open weights

Topics

inference, MoE, edge, serving, systems

Research notes

  • Method: Bandwidth-adaptive execution: full serving-stack co-design (model layout/loading, expert residency, CPU-GPU execution, agentic state reuse, runtime memory management) with continuous remapping of compute and state to observed hardware resources.
  • Key findings: Supports 20+ MoE models and real coding/tool-using agents from an 8GB laptop GPU to a single workstation GPU; Changes what local machines can serve: 35B model on a laptop, 284B on a gaming desktop, 753B GLM-5.2 on a single workstation GPU
  • Limitations: Evaluation is on the authors' hardware sweep; performance on other edge configurations and non-MoE architectures not characterized in the abstract.
  • Authors state the system is released at flashml.ai.