the.ai

Synthesis / Representation

verified

Speech Tokenization

Once Neural Audio Codec turns audio into discrete tokens, speech becomes a sequence-modelling problem and everything built for language applies to it. The tokens split into two kinds that do different jobs: semantic tokens carrying what was said, acoustic tokens carrying how it sounded.

Viz primitive · budget-splitacoustic-tokens = 18

acoustic-tokens holds 75% of the budget; rest holds the remaining 25%.

Acoustic tokens describing how it sounded against semantic tokens describing what was said, in tokens. Drag the acoustic levels up to watch them swamp the sequence — they already outnumber the semantic ones at any useful quality, which is why they are rarely generated one at a time.

18

Reviewed by opendroid · 2026-08-18