ASR Benchmarks 2026

Comprehensive word error rate (WER) comparison for modern speech recognition models

Last updated: January 18, 2026

This page provides up-to-date WER benchmarks for the most popular ASR models in 2026. All numbers are sourced from published papers, model cards, and verified community benchmarks. Use this as a reference when evaluating models for research or production deployment.

📊 How to Read These Tables

Lower WER is better. WER measures the percentage of words incorrectly transcribed. A WER of 2.0% means 98% of words are correct. Production systems typically target <5% WER for good user experience.

LibriSpeech Benchmarks

LibriSpeech is the standard academic benchmark for English ASR. It consists of clean audiobook recordings. Most research papers report results on test-clean and test-other subsets.

Model test-clean test-other Parameters Year Tags
Whisper Large-v3 1.4% 2.9% 1550M 2023 Production
Conformer-CTC (Google) 1.9% 3.9% 600M 2020 SOTA 2020
Wav2Vec 2.0 Large 1.9% 3.5% 317M 2020 Self-Supervised
HuBERT Large 1.9% 3.3% 317M 2021 Self-Supervised
Whisper Medium 2.4% 4.9% 769M 2022 Production
ContextNet (Google) 2.1% 4.6% 112M 2020 Streaming
Kaldi Chain (TDNN-F) 3.2% 7.6% ~20M 2019 Production
Whisper Small 3.0% 5.8% 244M 2022 Edge-Friendly
Whisper Base 4.8% 8.3% 74M 2022 Edge-Friendly

Common Voice Benchmarks

Common Voice represents more realistic, diverse speech from crowd-sourced recordings. Higher WER is expected due to accent variation and recording quality.

Model English (test) Spanish (test) German (test) Notes
Whisper Large-v3 5.2% 6.8% 7.1% Multilingual training
Wav2Vec 2.0 XLSR-53 7.3% 9.2% 8.6% 53 languages
Whisper Medium 6.8% 8.5% 9.2% Cost-effective choice
Kaldi (monolingual) 12.4% 15.8% 14.2% Requires language-specific tuning

TED-LIUM 3 Benchmarks

TED talks represent challenging spontaneous speech with varied topics, accents, and speaking styles.

Model WER (test) Real-Time Factor Hardware
Whisper Large-v3 3.8% 0.15x A100 GPU
Conformer-Transducer 4.2% 0.22x V100 GPU
Wav2Vec 2.0 Large 4.8% 0.18x V100 GPU
Kaldi Chain 6.3% 0.08x CPU (16 cores)
âš¡ Real-Time Factor Explained

RTF measures inference speed. 0.15x means processing 1 hour of audio takes 9 minutes. Lower is faster. Production systems typically need <0.3x for good UX.

Multilingual Performance

Cross-lingual ASR is critical for global products. These benchmarks show performance on 10 common languages.

Model Languages Avg WER Best For
Whisper Large-v3 99 7.2% General-purpose multilingual
Wav2Vec 2.0 XLSR-128 128 9.8% Low-resource languages
MMS (Meta) 1100+ 10.4% Rare language coverage
Google USM 100+ 8.6% Production streaming

Production Deployment Considerations

Benchmark WER doesn't always translate to production performance. Here's what matters for real-world deployments:

Factors Beyond WER

2026 Model Recommendations by Use Case

Use Case Recommended Model Why
Meeting Transcription Whisper Large-v3 Best accuracy, handles accents/noise
Real-Time Subtitles Conformer-RNN-T Streaming with low latency
Medical Transcription Fine-tuned Whisper Medium Cost-effective + customizable
Call Center Analytics Kaldi Chain CPU-efficient, proven at scale
On-Device (Mobile) Whisper Tiny + Quantization Small footprint, offline
Low-Resource Languages MMS or XLSR-128 1100+ language coverage

Working on ASR Research?

Labs and companies are hiring researchers who understand these benchmarks. Submit your profile to get matched with ASR research roles.

Get the weekly digest

Methodology & Sources

All benchmark numbers are sourced from:

Reporting Issues: If you find outdated or incorrect WER numbers, please email benchmarks@speechtechjobs.com with citations.

Related Resources