the.ai

Synthesis / Representation

verified

Speaker Embedding

A fixed-length vector that captures who is speaking rather than what they said. The same vector recognises a speaker across utterances and conditions a synthesiser to sound like them, which is why one representation serves verification and cloning at once — a duality worth noticing before building either.

Viz primitive · budget-splitenrolment-seconds = 4

enrolment-seconds holds 50% of the budget; rest holds the remaining 50%.

Seconds of enrolment audio against the single utterance being matched, in seconds. Drag the enrolment up to watch the estimate steady — a few seconds is enough, which is the whole security problem in one bar.

4

Reviewed by opendroid · 2026-08-18

  • arXiv:2301.02111 — Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers