the.ai

Synthesis / Generation

verified

Audio Generation

Generating music or sound effects rather than speech. The problem is structurally different: there is no transcript to constrain it, coherence is needed over minutes rather than seconds, and quality is judged by whether it is pleasant rather than whether it is correct.

Viz primitive · budget-splitstructure-tokens = 6

structure-tokens holds 25% of the budget; rest holds the remaining 75%.

Tokens spent on long-range structure against tokens spent on local detail, in tokens. Drag the structure up to watch coherence get the budget — it is what a listener notices and what likelihood scores least.

6

Reviewed by opendroid · 2026-08-18