A useful benchmark reports many metrics because different downstream decisions depend on different error properties. Our metric suite is organised into six groups; each answers a different question. All metrics operate on text that has been NFC-normalised, punctuation-stripped, whitespace-collapsed, and case-folded by a single versioned normaliser applied identically to reference and hypothesis.


Group 1 — Standard lexical accuracy

WER = (S + D + I) / N
CER = (S_c + D_c + I_c) / N_c
MER = (S + D + I) / (S + D + I + H)
WIL = 1 − H² / (N_ref × N_hyp)
WIP = 1 − WIL

Where S, D, I are substitutions, deletions, insertions; N is reference length; H is correctly aligned hits. Computed via jiwer.

<aside> 💡

Why it matters — CER is the primary metric for morphologically rich languages (Malayalam, Tamil, Telugu, Vietnamese) where word-level edit distance is disproportionately punishing. MER is bounded to [0, 1] even under heavy hallucination; WIL penalises missing and extra words symmetrically.

</aside>


Group 2 — Error-type decomposition

S_% = S / (H + S + D) × 100
D_% = D / (H + S + D) × 100
I_% = I / (H + S + D) × 100

The distribution across S, D, I is diagnostic:

No aggregate WER exposes this distinction.


Group 3 — Script-fair metrics

(our contribution)

These metrics exist to neutralise the script-mismatch effect in code-mixed speech. A reference that writes "telecom" in Bengali script and a hypothesis that writes it in Latin script are the same transcription; standard WER counts them as two errors.

Metric Definition Why it matters
lwWER WER after mapping every loanword to a canonical script in both reference and hypothesis Removes script-choice differences from the error count; the fairest single number for code-mixed content
toWER WER after transliterating both texts to ITRANS phonetic alphabet Phonetic upper bound; removes all script differences, not just mapped loanwords
OIWER WER after loanword normalisation and stripping all Latin-script tokens Native-language-only performance, free of code-switching influence
script_penalty WER − lwWER Direct measurement of how much of the headline WER is script mismatch rather than error

Table 3.1 — Four script-fair metrics. None of these are reported by any public ASR benchmark we are aware of.