The Humyn Score condenses a suite of speech-recognition accuracy metrics into one number between 0 and 1, where 1 is a perfect transcript. It is absolute: the score is anchored to measured error rates rather than to the other systems being compared, so a published score does not change when systems are added to or removed from the benchmark.

It is also coverage-adjusted. Systems support different languages, and those languages differ greatly in difficulty. Without adjustment, a system that declines the hardest languages scores higher than its capability warrants. The score corrects for this without penalizing a system for a language it does not claim to support.


1 · Metric selection

Seven accuracy metrics were available. Five are used. Two were removed because they restate information already present, which would otherwise let one measurement carry most of the score under two different names.

Metric Measures Weight — Indic Weight — International
Loanword-aware word error Word accuracy after borrowed vocabulary is scored correctly 50.0% 66.7%
Character error Sub-word accuracy; catches script and transliteration failure 20.0% 26.7%
Code-switch F1 Retention of embedded second-language vocabulary 15.0% —
Phoneme-informed error Error weighted by phonetic distance 10.0% —
Semantic similarity Meaning preservation 5.0% 6.7%
Plain word error Word accuracy, uncorrected 0.0% 0.0%
Word information lost Information-theoretic word accuracy 0.0% 0.0%

<aside> 💡

Why plain word error carries no weight. The loanword-aware variant is the same measurement with a correction applied; the two rank systems near-identically. Carrying both would spend most of the score on one quantity. The corrected form is kept because it is the fairer of the two — see section 5.

</aside>

<aside> 💡

Why word information lost carries no weight. Over 98% of its variance is recoverable from the other metrics. Including it adds weight to word accuracy under a different label rather than adding evidence.

</aside>

<aside> 💡

Why semantic similarity is capped at 5%. It is the only metric measuring a genuinely distinct property — roughly 70% of its variance is unique to it. But it discriminates poorly between systems, and it saturates: most transcripts score near the top of its range. It is retained as a floor detector for catastrophic output, not as a ranking signal.

</aside>

<aside> 💡

Why the secondary column drops two metrics. Code-switch F1 does not apply where there is no second-language convention to retain; measured variation between systems is negligible there. Phoneme-informed error was not reported for a fifth of that corpus, including two systems entirely — including it would have removed those systems from the comparison rather than scoring them. Remaining weights are renormalised to sum to 1.

</aside>


2 · Calculation

Each metric is converted to a quality value on 0–1, then combined by weight. Because the weights sum to 1 and each quality value is bounded, the result is bounded.

error metrics    q = 1 − min(value, 1)
quality metrics  q = value

Humyn Score = Σ weight_i × q_i

A score of 0.88 therefore means 12% weighted error. That reading is independent of which other systems appear in the table.


3 · Adjusting for unequal language coverage

This is the step that makes systems with different language support comparable. It proceeds in three stages:

  1. Aggregate per language. For each system and each language it supports, average every metric across that system's calls in that language.
  2. Separate system quality from language difficulty. Fit an additive model across all system-language cells: value ≈ baseline + system effect + language effect. Each system's effect is estimated only from its own observations.