the.ai

Multimodal / Generation

verified

Video Generation

Generating a video is not generating many images. Each frame has to be plausible and the sequence has to be consistent — the same person, the same room, objects that move the way objects move. A model that produces excellent frames independently produces a flickering mess, and fixing that is the entire problem.

Viz primitive · budget-splittemporal-attention = 20

temporal-attention holds 33% of the budget; rest holds the remaining 67%.

Compute spent relating frames to each other, against the compute of generating them independently, in equal units. Drag the temporal attention up to watch consistency take the budget — it is quadratic in duration, which is why clips are short and extended by conditioning.

20

Reviewed by opendroid · 2026-08-18

  • arXiv:2311.15127 — Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
  • arXiv:2304.08818 — Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models