the.ai

Multimodal / Foundations

verified

Audio-Visual Learning

Video comes with its own soundtrack, and the two are aligned for free — nobody had to label anything for the audio at a moment to correspond to the pixels at that moment. That correspondence is supervision lying around in enormous quantity, and it is what makes audio-visual pretraining possible without annotation.

Viz primitive · budget-splituninformative-audio = 20

uninformative-audio holds 40% of the budget; rest holds the remaining 60%.

Video whose audio says nothing about its pixels, against video where the correspondence carries signal, in hours. Drag the ambient share up to watch the useful supervision shrink — the pairing is free everywhere and informative in a fraction of it.

20

Reviewed by opendroid · 2026-08-18

  • arXiv:2104.11178 — VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
  • arXiv:2103.00020 — Learning Transferable Visual Models From Natural Language Supervision