the.ai

Synthesis / Methods

verified

Voice Conversion

Keep what was said and change who said it. That requires separating content from speaker identity in a representation where they are thoroughly entangled — the same phoneme sounds different in every voice, and the difference is precisely what you are trying to move.

Viz primitive · budget-splitbottleneck-width = 6

bottleneck-width holds 50% of the budget; rest holds the remaining 50%.

Capacity in the content bottleneck against the capacity a speaker embedding supplies, in equal units. Drag the bottleneck wider to watch identity leak back through the channel meant to carry only words.

6

Reviewed by opendroid · 2026-08-18

  • arXiv:2106.06103 — Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech