← Back to explorer

DSpark for DeepSeek-V4-Flash-0731 GGUFs

Type
other
Venue
Unsloth
Year
2026
Source
web
Access
free
Language
en
Added
2026-08-14T19:50:00Z
Verified
2026-08-14T19:50:00Z

Summary

Unsloth 2026-08-06 enablement. DeepSeek-V4-Flash-0731 is 284B/13B-active MoE, 1M context, QAT MXFP4 experts. Unsloth UD-Q8_K_XL is bit-exact vs official weights (KL~0, 100% top-token). DSpark (arXiv 2607.05147; llama.cpp PR 25784) is a draft-model speculative decoder claimed superior to naive MTP; Unsloth ships Q8_0 and BF16 drafter GGUFs. Tweet: ~1.4–2× faster, up to 120 tok/s, no accuracy change. Docs recommend --spec-type draft-dspark --spec-draft-n-max 3 (~1.9×); ~10 GB extra RAM. Auto-on in Unsloth Desktop. Also improved V4 chat jinja (reasoning_effort, tool-call reasoning_content) over 4000 conversations vs official baseline. Distinct from the base DeepSeek release.

Keywords

dspark · deepseek-v4-flash · unsloth · gguf · speculative-decoding · llama-cpp · x

Topics

speculative decoding, local GGUF inference, DeepSeek-V4

Research notes

  • Primary: Unsloth DeepSeek-V4 docs (DSpark section). Discord/X https://x.com/UnslothAI/status/2085368138393329703 via fxtwitter. GGUFs https://huggingface.co/unsloth/DeepSeek-V4-Flash-0731-GGUF (MIT). DSpark paper arXiv 2607.05147; llama.cpp PR 25784. Quant/docs item, not a hosted corpus, so no datasets_local row.