the.ai

Multimodal / Objectives

verified

Contrastive Learning

Instead of predicting a label, learn by telling matched pairs apart from mismatched ones. Pull the representations of things that belong together closer, push everything else away. No annotation is needed — the pairing itself is the supervision, and it can come from a caption, a crop, or the next second of audio.

Viz primitive · update-spectrumtemperature = 0.07

8 values. The left group decays steeply; the right group is 12% of the way to flat, and reads flatter than the left.

Similarity scores across a batch, positives against negatives. Drag the temperature to watch the objective sharpen or flatten.

0.07

Reviewed by opendroid · 2026-08-04

  • arXiv:1807.03748 — Representation Learning with Contrastive Predictive Coding
  • arXiv:2002.05709 — A Simple Framework for Contrastive Learning of Visual Representations