the.ai

Multimodal / Models

verified

Vision-Language Model

A language model that can see is mostly a language model. A vision encoder turns the image into vectors, a small bridge maps them into the language model's embedding space, and from there the image is just more tokens in the context. The heavy lifting stays where it already was.

Viz primitive · budget-splittrainable = 1

trainable holds 1% of the budget; rest holds the remaining 99%.

Trainable connector parameters against the frozen encoders on either side. Drag the bridge size to see how little is being learned.

1

Reviewed by opendroid · 2026-08-04