← Back to explorer

Phonon-2 (Fermion Research)

Type
model
Venue
Fermion Research
Year
2026
Source
project
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

Fermion Research releases Phonon-2, an English open speech-recognition model in a 164 MB download that it calls the most accurate open ASR model under 900 MB. Across the Open ASR Leaderboard's seven English sets it averages 5.21% word error, and every open model that scores better is at least 5.8 times its size. The model is distilled from NVIDIA's Parakeet TDT 0.6B v3 (2,508 MB) with encoder weights quantized to five learned levels (about 2.1 bits, base-3 packing of five digits per byte), holding the teacher's accuracy on LibriSpeech, beating it on AMI meetings and VoxPopuli parliamentary speech from a 15x smaller download. Speed: about 20 seconds per hour of audio on an Apple M5 MacBook Air via MLX (174x realtime), 142.8x realtime on eight Linux x86-64 cores, and up to 3,614x realtime batched on an A100. Released under CC-BY-4.0 with Hugging Face weights, a pip package (fermion-research), Docker images, and the Detta dictation app for Mac.

Keywords

ASR · speech recognition · quantization · distillation · Whisper · Parakeet · models · open weights

Topics

speech recognition, ASR, quantization, distillation

Research notes

  • Discovery: Manan Gupta (@yoitsmanan, verified) on 2026-09-29: https://x.com/yoitsmanan/status/2104990913886031993
  • Research page: https://fermionresearch.com/research/phonon-2/
  • Model: https://huggingface.co/FermionResearch/Phonon-2
  • pip install fermion-research; Docker: ghcr.io/fermionresearch/phonon-cpu
  • Connects to the collection's ASR, quantization, and model-compression entries.