The Humyn Score condenses a suite of speech-recognition accuracy metrics into one number between 0 and 1, where 1 is a perfect transcript. It is absolute: the score is anchored to measured error rates rather than to the other systems being compared, so a published score does not change when systems are added to or removed from the benchmark.
It is also coverage-adjusted. Systems support different languages, and those languages differ greatly in difficulty. Without adjustment, a system that declines the hardest languages scores higher than its capability warrants. The score corrects for this without penalizing a system for a language it does not claim to support.
Seven accuracy metrics were available. Five are used. Two were removed because they restate information already present, which would otherwise let one measurement carry most of the score under two different names.
| Metric | Measures | Weight — Indic | Weight — International |
|---|---|---|---|
| Loanword-aware word error | Word accuracy after borrowed vocabulary is scored correctly | 50.0% | 66.7% |
| Character error | Sub-word accuracy; catches script and transliteration failure | 20.0% | 26.7% |
| Code-switch F1 | Retention of embedded second-language vocabulary | 15.0% | — |
| Phoneme-informed error | Error weighted by phonetic distance | 10.0% | — |
| Semantic similarity | Meaning preservation | 5.0% | 6.7% |
| Plain word error | Word accuracy, uncorrected | 0.0% | 0.0% |
| Word information lost | Information-theoretic word accuracy | 0.0% | 0.0% |
<aside> 💡
Why plain word error carries no weight. The loanword-aware variant is the same measurement with a correction applied; the two rank systems near-identically. Carrying both would spend most of the score on one quantity. The corrected form is kept because it is the fairer of the two — see section 5.
</aside>
<aside> 💡
Why word information lost carries no weight. Over 98% of its variance is recoverable from the other metrics. Including it adds weight to word accuracy under a different label rather than adding evidence.
</aside>
<aside> 💡
Why semantic similarity is capped at 5%. It is the only metric measuring a genuinely distinct property — roughly 70% of its variance is unique to it. But it discriminates poorly between systems, and it saturates: most transcripts score near the top of its range. It is retained as a floor detector for catastrophic output, not as a ranking signal.
</aside>
<aside> 💡
Why the secondary column drops two metrics. Code-switch F1 does not apply where there is no second-language convention to retain; measured variation between systems is negligible there. Phoneme-informed error was not reported for a fifth of that corpus, including two systems entirely — including it would have removed those systems from the comparison rather than scoring them. Remaining weights are renormalised to sum to 1.
</aside>
Each metric is converted to a quality value on 0–1, then combined by weight. Because the weights sum to 1 and each quality value is bounded, the result is bounded.
error metrics q = 1 − min(value, 1)
quality metrics q = value
Humyn Score = Σ weight_i × q_i
A score of 0.88 therefore means 12% weighted error. That reading is independent of which other systems appear in the table.
This is the step that makes systems with different language support comparable. It proceeds in three stages: