← Back to explorer

Audio8-ASR-0.1B

Type
model
Venue
Audio8 / AutoArk-AI
Year
2026
Source
huggingface
Access
free
Language
en
Added
2026-08-14T19:50:00Z
Verified
2026-08-14T19:50:00Z

Summary

Audio8 (AutoArk) 0.1B-decoder ASR. LM 103.5M params; end-to-end unique params 324.0M. Qwen3-ASR audio encoder + MLP adapter + 8-layer Qwen-style causal LM. 16 kHz. Languages EN/ZH/FR/DE/JA/KO/Yue. Optional decode-time hotword boosting. Open ASR Leaderboard seven-split mean WER 7.03 / RTFx 741 on H200 (AMI 10.99, Earnings22 12.31, GigaSpeech 8.48, LibriSpeech clean/other 2.70/6.59, SPGISpeech 3.73, VoxPopuli 4.39). Internal WenetSpeech meeting/net CER 8.84/7.98. ONNX Runtime package ~1.1 GB peak; iOS ANE demo ~200 MB peak. Default examples cap audio at 30 s. Related paper Ark-ASR OPD (arXiv 2605.28139) is a 0.6B student distilled from Qwen-ASR on 100k hours; this 0.1B card is the compact public checkpoint.

Keywords

audio8 · asr · on-device · qwen3-asr · autoark · multilingual · x

Topics

on-device ASR, multilingual speech recognition, on-policy distillation

Research notes

  • Primary: HF model card (cc-by-nc-4.0). Discord/X https://x.com/SamuelZengML/status/2085321310067134527 via fxtwitter (Samuel Zeng / Audio8.ai). Paper arXiv 2605.28139 Data-Efficient On-Policy Distillation for Automatic Speech Recognition (Ark-ASR 0.6B, not this 0.1B checkpoint). Code https://github.com/AutoArk/open-audio-opd. Siblings Audio8-ASR-0.1B-onnx-runtime and Audio8-ASR-0.1B-iOS-ANE. Mirror AutoArk-AI/Audio8-ASR-0.1B. Open weights, not a new hosted corpus, so no datasets_local row.