- Status
- Verified
- Base model
- …
- Lineage
- …
- Unchanged from base
- …
- Trending
- #41
- Downloads, 30 days
- 1.3k
- Weights
- 164 MB
- Revision
Phonon-2 is the most accurate open speech recognition model for English under 900 MB. Across the Open ASR Leaderboard's seven English sets it averages 5.21 % word error, and every open model that scores better is at least 5.8 times its size.
At a glance
- Task
- Speech recognition
- Input
- audio
- Output
- text
- Architecture
- Parakeet Tdt Five Value
- Library
- mlx
- License
- CC BY 4.0Commercial use
- Languages
- en
- Base model
- Quantized from nvidia/parakeet-tdt-0.6b-v3
- Released
- Sep 2026
- Updated
- Sep 2026
- Likes
- 28
- Downloads, all time
- 170
Benchmarks
8 results, self reported by the authors.
| Benchmark | Score | Date | Evidence |
|---|---|---|---|
| hf-audio/open-asr-leaderboardmean_wer | 5.21 | Sep 2026 | Reported by source |
| hf-audio/open-asr-leaderboardlibrispeech_clean_wer | 1.72 | Sep 2026 | Reported by source |
| hf-audio/open-asr-leaderboardlibrispeech_other_wer | 3.92 | Sep 2026 | Reported by source |
| hf-audio/open-asr-leaderboardami_wer | 9.37 | Sep 2026 | Reported by source |
| hf-audio/open-asr-leaderboardearnings22_wer | 6.96 | Sep 2026 | Reported by source |
| hf-audio/open-asr-leaderboardgigaspeech_wer | 8.35 | Sep 2026 | Reported by source |
| hf-audio/open-asr-leaderboardspgispeech_wer | 3.7 | Sep 2026 | Reported by source |
| hf-audio/open-asr-leaderboardvoxpopuli_wer | 2.46 | Sep 2026 | Reported by source |
Run it
Pinned to the indexed revision.
hf download FermionResearch/Phonon-2 --revision e357655f6325aa70d4800d125273a5ed3a703b9cThrough the hub: the same tools, each file from a source that is up (Hugging Face, ModelScope, IPFS), at this revision. The second line checks every file against its address.
export HF_ENDPOINT=https://gethologram.ai
cd "$(hf download FermionResearch/Phonon-2 --quiet)" && curl -s $HF_ENDPOINT/FermionResearch/Phonon-2/resolve/main/SHA256SUMS | sha256sum -c --quietRead the full model card
Phonon-2
Phonon-2 is the most accurate open speech recognition model for English under 900 MB. Across the Open ASR Leaderboard's seven English sets it averages 5.21 % word error, and every open model that scores better is at least 5.8 times its size. Set for set it holds the accuracy of its 2.5 GB full-precision teacher, reaching 100.8 % of the teacher's word accuracy on parliamentary speech and beating it on meetings, from a download 15 times smaller. Its encoder holds each weight at one of five learned levels in about 2.1 bits.
It transcribes an hour of audio in about 20 seconds on an M5 MacBook Air (174x realtime), at 143x on eight Zen 5 cores (16 vCPU) and at 6,680x on one H100 in batches of 128.
Benchmarks
| Model | Download | LS clean | LS other | AMI | Earnings-22 | GigaSpeech | SPGISpeech | VoxPopuli | Average |
|---|---|---|---|---|---|---|---|---|---|
| Phonon-2 | 164 MB | 1.72 | 3.92 | 9.37 | 6.96 | 8.35 | 3.70 | 2.46 | 5.21 |
| Parakeet TDT 0.6B v3, teacher† | 2,508 MB | 1.52 | 3.13 | 9.42 | 5.85 | 7.99 | 3.63 | 3.19 | 4.96 |
| Parakeet Redux | 178 MB | 1.94 | 4.35 | 9.16 | 7.90 | 8.62 | 4.01 | 3.87 | 5.69 |
| Phonon-1 | 415 MB | 2.11 | 5.03 | 10.31 | 12.34 | 8.73 | 3.67 | 3.73 | 6.56 |
| Canary 180M Flash† | 737 MB | 1.52 | 3.42 | 12.09 | 8.33 | 8.87 | 2.04 | 3.57 | 5.69 |
| Voxtral Mini 4B Realtime† | ≈8,000 MB* | 1.62 | 4.94 | 13.34 | 9.31 | 8.80 | 2.23 | 2.60 | 6.12 |
| Whisper large-v3-turbo† | 1,618 MB | 2.13 | 3.71 | 13.88 | 8.09 | 8.47 | 2.79 | 7.02 | 6.58 |
| Nemotron 3.5 ASR Streaming 0.6B† | 2,368 MB | 2.83 | 6.79 | 13.43 | 15.30 | 9.86 | 3.27 | 4.24 | 7.96 |
† Open ASR Leaderboard's published row; the other rows use its code on the full test sets. * Size from the parameter count at 16 bits.
Run it
Phonon-2 is the model inside Detta, the dictation app for the Mac.
On Apple silicon, from the command line:
pip install fermion-research
pip install mlx mlx-audio mlx-lm soundfile scipy zstandard
fermion transcribe recording.wav
The same package runs on Linux (x86-64 and Arm) and Windows CPUs; the engines are at github.com/fermionresearch/phonon. With Docker, on a CPU or a GPU:
docker run --rm -v "$PWD":/audio ghcr.io/fermionresearch/phonon-cpu:2.0.2 transcribe /audio/recording.wav --model phonon-2
docker run --rm --gpus all -v "$PWD":/audio ghcr.io/fermionresearch/phonon-cuda:1.0.3 transcribe /audio/recording.wav --model phonon-2
Notes
Based on parakeet-tdt-0.6b-v3 by NVIDIA; the tokenizer and output conventions (punctuation, casing, numerals) are the
original's. Licence CC-BY-4.0, same as the original; NOTICE lists the changes. The command line and this
repository's code are Apache 2.0.
Derived on Sep 30, 2026 from Hugging Face at revision e357655f, README.md , config.json .
12 files, 164 MB. Every download is checked against its address.