Jump to Benchmark:
This page provides up-to-date WER benchmarks for the most popular ASR models in 2026. All numbers are sourced from published papers, model cards, and verified community benchmarks. Use this as a reference when evaluating models for research or production deployment.
Lower WER is better. WER measures the percentage of words incorrectly transcribed. A WER of 2.0% means 98% of words are correct. Production systems typically target <5% WER for good user experience.
LibriSpeech Benchmarks
LibriSpeech is the standard academic benchmark for English ASR. It consists of clean audiobook recordings. Most research papers report results on test-clean and test-other subsets.
| Model | test-clean | test-other | Parameters | Year | Tags |
|---|---|---|---|---|---|
| Whisper Large-v3 | 1.4% | 2.9% | 1550M | 2023 | Production |
| Conformer-CTC (Google) | 1.9% | 3.9% | 600M | 2020 | SOTA 2020 |
| Wav2Vec 2.0 Large | 1.9% | 3.5% | 317M | 2020 | Self-Supervised |
| HuBERT Large | 1.9% | 3.3% | 317M | 2021 | Self-Supervised |
| Whisper Medium | 2.4% | 4.9% | 769M | 2022 | Production |
| ContextNet (Google) | 2.1% | 4.6% | 112M | 2020 | Streaming |
| Kaldi Chain (TDNN-F) | 3.2% | 7.6% | ~20M | 2019 | Production |
| Whisper Small | 3.0% | 5.8% | 244M | 2022 | Edge-Friendly |
| Whisper Base | 4.8% | 8.3% | 74M | 2022 | Edge-Friendly |
Common Voice Benchmarks
Common Voice represents more realistic, diverse speech from crowd-sourced recordings. Higher WER is expected due to accent variation and recording quality.
| Model | English (test) | Spanish (test) | German (test) | Notes |
|---|---|---|---|---|
| Whisper Large-v3 | 5.2% | 6.8% | 7.1% | Multilingual training |
| Wav2Vec 2.0 XLSR-53 | 7.3% | 9.2% | 8.6% | 53 languages |
| Whisper Medium | 6.8% | 8.5% | 9.2% | Cost-effective choice |
| Kaldi (monolingual) | 12.4% | 15.8% | 14.2% | Requires language-specific tuning |
TED-LIUM 3 Benchmarks
TED talks represent challenging spontaneous speech with varied topics, accents, and speaking styles.
| Model | WER (test) | Real-Time Factor | Hardware |
|---|---|---|---|
| Whisper Large-v3 | 3.8% | 0.15x | A100 GPU |
| Conformer-Transducer | 4.2% | 0.22x | V100 GPU |
| Wav2Vec 2.0 Large | 4.8% | 0.18x | V100 GPU |
| Kaldi Chain | 6.3% | 0.08x | CPU (16 cores) |
RTF measures inference speed. 0.15x means processing 1 hour of audio takes 9 minutes. Lower is faster. Production systems typically need <0.3x for good UX.
Multilingual Performance
Cross-lingual ASR is critical for global products. These benchmarks show performance on 10 common languages.
| Model | Languages | Avg WER | Best For |
|---|---|---|---|
| Whisper Large-v3 | 99 | 7.2% | General-purpose multilingual |
| Wav2Vec 2.0 XLSR-128 | 128 | 9.8% | Low-resource languages |
| MMS (Meta) | 1100+ | 10.4% | Rare language coverage |
| Google USM | 100+ | 8.6% | Production streaming |
Production Deployment Considerations
Benchmark WER doesn't always translate to production performance. Here's what matters for real-world deployments:
Factors Beyond WER
- Latency: Whisper is offline-only. For streaming, use Conformer-RNN-T or Kaldi.
- Domain Adaptation: Fine-tuning Whisper on 10-100 hours of custom data can reduce WER by 20-40%.
- Cost: Whisper Large costs ~$0.02/minute on A100. Kaldi can run <$0.001/minute on CPU.
- Robustness: Models trained on clean data (LibriSpeech) often fail on noisy production audio.
- Language Support: Whisper handles 99 languages out-of-the-box. Kaldi requires separate models per language.
2026 Model Recommendations by Use Case
| Use Case | Recommended Model | Why |
|---|---|---|
| Meeting Transcription | Whisper Large-v3 | Best accuracy, handles accents/noise |
| Real-Time Subtitles | Conformer-RNN-T | Streaming with low latency |
| Medical Transcription | Fine-tuned Whisper Medium | Cost-effective + customizable |
| Call Center Analytics | Kaldi Chain | CPU-efficient, proven at scale |
| On-Device (Mobile) | Whisper Tiny + Quantization | Small footprint, offline |
| Low-Resource Languages | MMS or XLSR-128 | 1100+ language coverage |
Working on ASR Research?
Labs and companies are hiring researchers who understand these benchmarks. Submit your profile to get matched with ASR research roles.
Get the weekly digestMethodology & Sources
All benchmark numbers are sourced from:
- Official model papers and Hugging Face model cards
- Papers With Code leaderboards (verified)
- SpeechBrain and ESPnet published recipes
- Community benchmarks (GitHub issues, verified reproductions)
Reporting Issues: If you find outdated or incorrect WER numbers, please email benchmarks@speechtechjobs.com with citations.