← Back to explorer

fuse-1 Lite

Type
model
Venue
Akahsizrr (Hugging Face)
Year
2026
Source
huggingface
Access
free
Language
en
Added
2026-08-14T18:32:00Z
Verified
2026-08-14T18:32:00Z

Summary

fuse-1 Lite (Fuse3ForCausalLM) freezes LiquidAI/LFM2.5-2.6B (2.70B) and 960 coding experts taken from Qwen/Qwen3.6-35B-A3B (3.02B), training only ~2.0M router+scale params. 30 host layers, 32 experts/layer, top-8; after 300 steps (~8.3 min on a Modal L4, 55 examples) the router uses 11/30 layers and zeros the rest. Card: 5.72B total, Apache-2.0, trust_remote_code. HumanEval pass@1 listed TBD. Siblings 4-bit/8-bit/MLX/GGUF/vLLM; the limitations section still warns the 55-example router may not generalize and that some backends need custom code.

Keywords

fuse-1-lite · lfm2 · qwen3.6 · moe · expert-transplant · coding · huggingface

Topics

model fusion, mixture-of-experts, code LLMs

Research notes

  • Primary: HF model card (apache-2.0, transformers, pipeline_tag text-generation, downloads=2223 likes=118 checked 2026-08-14). Citation author Vasko Djack. Base models LiquidAI/LFM2.5-2.6B and Qwen/Qwen3.6-35B-A3B. Discord posted the HF model pointer. Weights/checkpoint item, not a new hosted pretraining corpus, so no datasets_local row. Quantized siblings (fuse-1-Lite-4bit/8bit/MLX/GGUF/vLLM) not copied into hf_* fields.