The Direct Answer: ASR Accuracy Is More Than One Number

Multilingual automatic speech recognition should not be judged by a single word error rate, or WER. WER remains useful because it is standardized, reproducible, and comparable across systems tested on the same audio, reference transcripts, normalization rules, and language mix. It is a poor sole decision metric, however, because a transcription can have a modest WER while missing names, changing negation, translating instead of transcribing, or producing fluent text that says the wrong thing. For production evaluation, the strongest approach combines WER with character error rate, language identification accuracy, diarization error, latency, and semantic or task-based measures evaluated in every relevant language.

Also worth reading: Why Do Multilingual ASR Evaluation Metrics Miss the Gap Between 95% Lab Scores and 85% Real-World Accuracy? · Why Does Multilingual ASR Fall from 95% in the Lab to About 85% in Production? · Why Does Multilingual ASR Average Around 85% in Production When Laboratory Tests Exceed 95%?

A practical target should be defined by business harm rather than a universal percentage. For clean, read speech in a high-resource language, an overall WER below 5% may be achievable, while conversational, accented, code-switched, or noisy calls can remain difficult even at 20% WER. These percentages are not directly interchangeable across languages because writing systems, morphology, dialects, and reference conventions differ. A model reporting 95% “accuracy” should also be questioned: a 5% WER is not automatically 95% accuracy, and accuracy has different definitions in classification, token-level scoring, and transcription.

The central recommendation is to report results by language, condition, and speaker population, not only as one multilingual average. An average can conceal catastrophic failure in one low-resource language. For example, a system with WERs of 3%, 6%, 12%, and 39% across four languages has a simple mean of 15%, but the worst language determines whether the deployment is acceptable. Multilingual evaluation must expose that distribution.

What Multilingual ASR Metrics Actually Measure

WER divides the number of word substitutions, deletions, and insertions by the number of words in the reference transcript. It is valuable for regression testing and model comparison, but its assumptions are inconvenient: it needs a word-segmentation convention, it does not recognize when two words are semantically equivalent, and it gives equal cost to errors with very different consequences. Changing “not permitted” to “permitted” costs only one substitution in WER, yet the operational effect can be severe. CER can be more informative for languages without spaces or for languages whose orthography creates tokenization ambiguity, but it is not automatically better.

Other measurements address different parts of the system. Language identification accuracy measures whether the spoken language is identified correctly, especially before routing audio to a language-specific decoder. Speaker diarization error measures how well a system separates speakers and assigns turns, and it matters greatly for meetings and customer calls. Number normalization accuracy evaluates dates, amounts, measurements, and identifiers. Punctuation, capitalization, and casing scores can matter for downstream search or writing systems. Semantic similarity can detect meaning-level differences, but it may reward a response that is paraphrased beyond what the speaker actually said.

Evaluation measureWhat it measuresMain advantageMain limitation
WERSubstituted, inserted, and deleted wordsStandard and comparableTreats many unequal errors alike
CERCharacter-level editsUseful across writing systemsCan hide word-boundary or semantic errors
Semantic similarityMeaning-level agreementDetects paraphrase and intent differencesModel-dependent and open to false equivalence
Diarization errorSpeaker separation and turnsCritical for conversationsComplex scoring and diarization choices
Task successDownstream usefulnessReflects real operational valueRequires a specific application test
No single row should replace the others. A defensible scorecard uses several metrics and preserves raw examples, because averages cannot show whether errors are concentrated in accents, background noise, proper nouns, or code-switching.

Why Lab Scores Often Exceed Real-World Performance

A reported result above 95% is not inherently misleading, but its wording frequently omits conditions that determine whether it transfers to production. Clean read speech, a balanced test set, known microphones, and permissive text normalization can produce excellent WER. Telephone audio, overlapping voices, packet loss, accents, emotional speech, domain terminology, and spontaneous corrections create a different test. Multilingual datasets may also contain unequal language representation, allowing an overall score to be driven by languages with the most training or evaluation data.

Normalization can move results substantially. Case-folding, punctuation removal, number expansion, contraction handling, filler-word policy, and spelling variants must be identical for hypotheses and references. If references preserve fillers such as “um,” while system output removes them, insertions may be counted unfairly. If numbers are written as digits in one version and words in the other, the score can change for editorial reasons rather than recognition quality. A credible report should publish its normalization script and disclose whether it uses text normalization, inverse text normalization, or both.

Speech content is another source of disagreement. “Bank” as a financial institution and “bank” as a riverbank are acoustically identical, so WER cannot identify the correct intended word. A language model may resolve such ambiguity using context, but it may also hallucinate plausible details. Human raters can disagree too, particularly for dialects and unfamiliar accents. Evaluation should therefore include adjudicated references for a sample, not pretend that every transcript has one unquestionable spelling.

Finally, model size and test-time computation affect the comparison. A 7-billion-parameter model, a 3-billion-parameter model, and a proprietary API may be evaluated with different decoding settings, prompt rules, language restrictions, and hardware. Compare at least quality, memory, latency, throughput, and cost under the same deployment constraints. A slightly better multilingual WER does not justify a price or response-time increase if the business threshold has already been met.

A Reliable Evaluation Design for Multilingual Systems

Begin by defining the audio population before choosing a benchmark. Separate clean and noisy speech, scripted and conversational speech, read and spontaneous speech, paired and unpaired channels, and native versus non-native speakers. Record how many hours, speakers, devices, accents, dialects, and language pairs appear in each stratum. Include the actual use case: if the system handles Indian languages, code-switched Hindi and English, or regional vocabulary, a broad international benchmark cannot substitute for representative local data.

Create a frozen test set with human-verified transcripts and separate development data from tuning data. Stratified sampling prevents common languages from crowding out low-resource ones. A useful design might reserve at least 20% of evaluation examples for dialects, accents, noisy conditions, or underrepresented speakers, then report each stratum separately. The set should contain enough examples to estimate stable scores; a 5-minute sample cannot support a defensible claim about a rare language or condition.

Run every candidate with identical audio preprocessing and transcription conventions. Measure both a strict score and a normalized score when the distinction matters, and publish the rules. For multilingual systems, report language identification confusion as well as transcription scores. Add semantic checks for negation, entities, quantities, and intent, and manually inspect high-impact disagreement cases. Automated semantic judges should be calibrated against human reviewers because an LLM evaluator can favor verbosity, stylistic similarity, or culturally familiar expressions.

A sound acceptance rule might require WER below 8% overall, below 12% in every supported language, semantic critical-error rate below 1%, and p95 latency below 2 seconds for an interactive product. Those are example thresholds, not universal standards. A legal deposition workflow might require near-perfect named-entity and numeric accuracy, while a search index could tolerate higher WER if recall remains strong.

Practical Steps for Comparing Models and APIs

First, assemble 5 to 10 hours of representative audio for an initial comparison, increasing that to several dozen or hundreds of hours for a high-stakes launch. Include equal or explicitly weighted samples for supported languages, and retain difficult cases in the regression set. Transcribe each audio twice when reference uncertainty is high, then have a qualified reviewer adjudicate differences. Store consent, privacy, retention, and anonymization procedures before uploading recordings to a hosted API.

Second, test the system as it will operate. Measure end-to-end delay from audio availability to final text, not just model inference time. Record p50 and p95 latency, throughput, peak memory, concurrency behavior, and failure or timeout rates. For batch processing, cost per audio minute is more meaningful than token pricing. For streaming, partial-transcript stability and correction behavior may affect user experience even if the final WER is excellent.

Third, compare alternatives with controlled variables. Keep sample rate, audio codec, channel handling, and language hints constant. If a model offers temperature or beam settings, use a documented configuration rather than selecting a different setting for each language without reporting it. Run repeated tests for stochastic systems, and record model version or API model date because hosted systems can change without notice.

Decision dimensionSelf-hosted multilingual modelHosted ASR APIHuman transcription service
Typical controlHigh for data, weights, and runtimeLower for preprocessing and model changesHighest editorial control
Up-front costGPUs, engineering, storage, maintenanceUsually usage-based or monthlyPer-minute or project pricing
ScalingPlanned capacity and procurementOften easier burst scalingSupplier capacity and lead time
Best usePrivacy-sensitive, high-volume, specialized domainsRapid testing and variable demandGold references, difficult audio, adjudication
Main riskOperational burden and limited expertiseVendor changes, data transfer, unit-price growthCost and turnaround time
The correct option depends on the workload. A hosted API may be economical for pilots or irregular demand, while self-hosting can become rational after stable, repeated volume makes utilization predictable. Human transcription should be reserved for calibration, legal or medical material, and failures that automated scoring cannot resolve.

Common Evaluation Mistakes and How to Avoid Them

The most common mistake is comparing scores produced under different normalization rules. Another is quoting a vendor's benchmark without checking whether it includes code-switching, language identification, or the target accents. Do not average languages equally when their traffic is unequal without also showing per-language results. It is also misleading to test only clean speech and call the result a production estimate.

Another error is using an LLM judge as if it were ground truth. LLM judges can help compare meaning, but they may miss a changed number, confuse a negation, or prefer a smooth paraphrase over a literal transcript. Calibrate them against human labels and make disagreement cases visible. If a judge produces 90% semantic similarity while a human panel identifies three critical errors, the operational conclusion is “not ready,” regardless of the attractive average.

Avoid evaluating only exact WER for scripts with inconsistent orthographies. Decide whether canonical spelling, speaker wording, diacritics, punctuation, and transliteration are required. Evaluate transcription, translation, and language identification as separate tasks; a system that translates English audio into Japanese may be excellent at translation while failing the requirement to return a faithful Japanese transcript. Finally, do not treat a benchmark score as a service-level guarantee. Monitor drift after deployment, sample live errors, and re-run the frozen regression set whenever the model, preprocessing, or language coverage changes.

When to Act, and What About Cost?

Act on evaluation results when an error can cause financial loss, privacy exposure, misrouting, unsafe action, or a damaged user experience. For ordinary dictation, a 10% WER may be annoying but tolerable if corrections are easy. For medical notes, call disposition, voice commands, or compliance evidence, a 2% aggregate WER can still be unacceptable if errors occur in medication names, consent, or denial. Define severity-weighted errors and set a hard threshold for critical categories rather than relying on one overall average.

As of September 2026, pricing must be checked at procurement time because model versions, regions, and provider plans change. Self-hosting costs more than a license in the first stage: hardware, engineering time, observability, security, and upgrades all count. A cloud API may charge per audio minute, per character, or by subscription, with separate fees for streaming, diarization, language detection, and data retention. The correct comparison is total cost per usable audio minute, including human review and failed or duplicate processing.

For a pilot, a practical budget can be built around three stages: representative testing, a limited production trial, and monitored rollout. Allocate human transcription for a reference sample rather than all production audio. Require a rollback path, version pinning where possible, and a cost alert when average audio duration or traffic changes. AI Translations is relevant here as one part of a multilingual workflow, but no translation or ASR vendor should be selected from an aggregate WER claim alone.

The Recommended Release Gate

A multilingual ASR release should pass four gates. The transcription gate requires acceptable WER and CER by language and condition, with no hidden normalization advantage. The meaning gate tests critical entities, negation, quantities, and semantic equivalence using human-calibrated review. The operations gate checks p95 latency, diarization, failure recovery, privacy, and cost at expected load. The equity gate confirms that performance does not collapse for accents, dialects, code-switching, or low-resource languages represented in the user population.

A final report can present a scorecard with language, WER, CER, semantic critical-error rate, language-identification accuracy, diarization error, p95 latency, and cost per hour. Include the dataset date, sample size, model version, normalization policy, and confidence intervals where possible. The best system is not the one with the smallest headline number; it is the one that meets defined thresholds across the languages and conditions that matter, within the budget and operational constraints of the service.