the.ai

Synthesis / Evaluation

verified

Synthesis Evaluation

There is no word error rate for how good speech sounds. The field's standard is a mean opinion score — people rating samples one to five — which is slow, expensive, and not comparable between papers because the raters and instructions differ every time.

Viz primitive · budget-splithuman-ratings = 4

human-ratings holds 13% of the budget; rest holds the remaining 87%.

Judgements from human raters against judgements from automatic proxies, in ratings. Drag the human share up to watch the measurement become trustworthy and expensive at the same rate.

4

Reviewed by opendroid · 2026-08-18

  • arXiv:2304.09116 — NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers