Phonon-2 (Fermion Research)
- Type
- model
- Venue
- Fermion Research
- Year
- 2026
- Source
- project
- Access
- public
- Language
- en
- Added
- 2026-09-30
- Verified
- 2026-09-30
Summary
Fermion Research releases Phonon-2, an English open speech-recognition model in a 164 MB download that it calls the most accurate open ASR model under 900 MB. Across the Open ASR Leaderboard's seven English sets it averages 5.21% word error, and every open model that scores better is at least 5.8 times its size. The model is distilled from NVIDIA's Parakeet TDT 0.6B v3 (2,508 MB) with encoder weights quantized to five learned levels (about 2.1 bits, base-3 packing of five digits per byte), holding the teacher's accuracy on LibriSpeech, beating it on AMI meetings and VoxPopuli parliamentary speech from a 15x smaller download. Speed: about 20 seconds per hour of audio on an Apple M5 MacBook Air via MLX (174x realtime), 142.8x realtime on eight Linux x86-64 cores, and up to 3,614x realtime batched on an A100. Released under CC-BY-4.0 with Hugging Face weights, a pip package (fermion-research), Docker images, and the Detta dictation app for Mac.
Keywords
ASR · speech recognition · quantization · distillation · Whisper · Parakeet · models · open weights
Topics
speech recognition, ASR, quantization, distillation
Research notes
- Discovery: Manan Gupta (@yoitsmanan, verified) on 2026-09-29: https://x.com/yoitsmanan/status/2104990913886031993
- Research page: https://fermionresearch.com/research/phonon-2/
- Model: https://huggingface.co/FermionResearch/Phonon-2
- pip install fermion-research; Docker: ghcr.io/fermionresearch/phonon-cpu
- Connects to the collection's ASR, quantization, and model-compression entries.