Synthesis / Evaluation
verifiedSynthesis Evaluation
There is no word error rate for how good speech sounds. The field's standard is a mean opinion score — people rating samples one to five — which is slow, expensive, and not comparable between papers because the raters and instructions differ every time.
Automatic proxies exist and are used because they are cheap: predicted opinion scores, speaker similarity via embeddings, and intelligibility measured by running an ASR system over the output and reporting word error rate. That last one is genuinely useful and genuinely narrow — a synthesiser can be perfectly intelligible and sound wrong, and a WER of near zero says nothing about prosody.
Opinion scores are ordinal, so averaging them is already a modelling choice, and the interval around the mean has a rater component as well as a sample one — twenty raters on ten samples is a different measurement from ten raters on twenty. Papers reporting a mean without the rater count and the interval have reported one number and no way to compare it.
human-ratings holds 13% of the budget; rest holds the remaining 87%.
Judgements from human raters against judgements from automatic proxies, in ratings. Drag the human share up to watch the measurement become trustworthy and expensive at the same rate.
Reviewed by opendroid · 2026-08-18
- arXiv:2304.09116 — NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers