A useful benchmark reports many metrics because different downstream decisions depend on different error properties. Our metric suite is organised into six groups; each answers a different question. All metrics operate on text that has been NFC-normalised, punctuation-stripped, whitespace-collapsed, and case-folded by a single versioned normaliser applied identically to reference and hypothesis.
WER = (S + D + I) / N
CER = (S_c + D_c + I_c) / N_c
MER = (S + D + I) / (S + D + I + H)
WIL = 1 − H² / (N_ref × N_hyp)
WIP = 1 − WIL
Where S, D, I are substitutions, deletions, insertions; N is reference length; H is correctly aligned hits. Computed via jiwer.
<aside> 💡
Why it matters — CER is the primary metric for morphologically rich languages (Malayalam, Tamil, Telugu, Vietnamese) where word-level edit distance is disproportionately punishing. MER is bounded to [0, 1] even under heavy hallucination; WIL penalises missing and extra words symmetrically.
</aside>
S_% = S / (H + S + D) × 100
D_% = D / (H + S + D) × 100
I_% = I / (H + S + D) × 100
The distribution across S, D, I is diagnostic:
No aggregate WER exposes this distinction.
(our contribution)
These metrics exist to neutralise the script-mismatch effect in code-mixed speech. A reference that writes "telecom" in Bengali script and a hypothesis that writes it in Latin script are the same transcription; standard WER counts them as two errors.
| Metric | Definition | Why it matters |
|---|---|---|
| lwWER | WER after mapping every loanword to a canonical script in both reference and hypothesis | Removes script-choice differences from the error count; the fairest single number for code-mixed content |
| toWER | WER after transliterating both texts to ITRANS phonetic alphabet | Phonetic upper bound; removes all script differences, not just mapped loanwords |
| OIWER | WER after loanword normalisation and stripping all Latin-script tokens | Native-language-only performance, free of code-switching influence |
| script_penalty | WER − lwWER | Direct measurement of how much of the headline WER is script mismatch rather than error |
Table 3.1 — Four script-fair metrics. None of these are reported by any public ASR benchmark we are aware of.