the.ai

Synthesis / Architecture

verified

Vocoder

A spectrogram has thrown away the phase, so turning one back into sound is not an inverse transform — it is a generation problem. The vocoder is the model that does it, and for years it was the bottleneck: the acoustic model produced a plausible spectrogram and the vocoder made it sound synthetic.

Viz primitive · budget-splitparallel-samples = 6

parallel-samples holds 50% of the budget; rest holds the remaining 50%.

Samples generated in parallel against samples that must wait for the one before, in samples. Drag the parallelism up to watch the sequential remainder vanish — at the far left this is one sample at a time, at 24 kHz.

6

Reviewed by opendroid · 2026-08-18

  • arXiv:2206.04658 — BigVGAN: A Universal Neural Vocoder with Large-Scale Training
  • arXiv:2010.05646 — HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis