the.ai

Vision / Dense Prediction

verified

Instance Segmentation

Semantic segmentation says these pixels are all 'person'. Instance segmentation says these are person one and those are person two, which is the difference between a picture that can be described and one that can be counted. Every application that needs a number — cells in a slide, items on a belt, people in a crowd — needs this and not the other.

Viz primitive · budget-splitmask-heads = 6

mask-heads holds 13% of the budget; rest holds the remaining 87%.

Compute in the per-instance mask heads, against the backbone that runs once for the whole image, in equal units. Drag the object count up to watch the per-object cost overtake the per-image one — the same model, two different cost models, decided by the photograph.

6

Reviewed by opendroid · 2026-08-18