the.ai

Speech / Synthesis

verified

Text to Speech

Reading a sentence aloud is not one problem but two: deciding how it should sound, and producing a waveform that sounds that way. Neural vocoders largely settled the second. The first is harder than it looks, because a sentence has many correct readings, and a model trained to minimise error will average them into a reading nobody would give.

Viz primitive · update-spectrumregression-weight = 0.2 · bars = 16

16 values. The left group decays steeply; the right group is 36% of the way to flat, and reads flatter than the left.

Harmonic detail in one frame, as recorded and as a mean-seeking loss predicts it. Drag the regression weight up to watch the peaks collapse toward their own average — the mean of every plausible reading is a reading that sounds muffled.

0.2

Reviewed by opendroid · 2026-08-18

  • arXiv:1712.05884 — Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions
  • arXiv:2010.05646 — HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis