Audio8-ASR-0.1B
- Type
- model
- Venue
- Audio8 / AutoArk-AI
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en
- Added
- 2026-08-14T19:50:00Z
- Verified
- 2026-08-14T19:50:00Z
Summary
Audio8 (AutoArk) 0.1B-decoder ASR. LM 103.5M params; end-to-end unique params 324.0M. Qwen3-ASR audio encoder + MLP adapter + 8-layer Qwen-style causal LM. 16 kHz. Languages EN/ZH/FR/DE/JA/KO/Yue. Optional decode-time hotword boosting. Open ASR Leaderboard seven-split mean WER 7.03 / RTFx 741 on H200 (AMI 10.99, Earnings22 12.31, GigaSpeech 8.48, LibriSpeech clean/other 2.70/6.59, SPGISpeech 3.73, VoxPopuli 4.39). Internal WenetSpeech meeting/net CER 8.84/7.98. ONNX Runtime package ~1.1 GB peak; iOS ANE demo ~200 MB peak. Default examples cap audio at 30 s. Related paper Ark-ASR OPD (arXiv 2605.28139) is a 0.6B student distilled from Qwen-ASR on 100k hours; this 0.1B card is the compact public checkpoint.
Keywords
audio8 · asr · on-device · qwen3-asr · autoark · multilingual · x
Topics
on-device ASR, multilingual speech recognition, on-policy distillation
Research notes
- Primary: HF model card (cc-by-nc-4.0). Discord/X https://x.com/SamuelZengML/status/2085321310067134527 via fxtwitter (Samuel Zeng / Audio8.ai). Paper arXiv 2605.28139 Data-Efficient On-Policy Distillation for Automatic Speech Recognition (Ark-ASR 0.6B, not this 0.1B checkpoint). Code https://github.com/AutoArk/open-audio-opd. Siblings Audio8-ASR-0.1B-onnx-runtime and Audio8-ASR-0.1B-iOS-ANE. Mirror AutoArk-AI/Audio8-ASR-0.1B. Open weights, not a new hosted corpus, so no datasets_local row.