the.ai

Adversarial / Evaluation

verified

Red Teaming

Attack your own system before someone else does. For a language model that means deliberately searching for prompts that produce the behaviour you claim it will not produce, and treating each success as a finding rather than an embarrassment. Done manually it does not scale; the interesting version uses one model to attack another.

Viz primitive · budget-splitcategories-probed = 6

categories-probed holds 25% of the budget; rest holds the remaining 75%.

Harm categories the exercise actually probed against those left untouched, in categories. Drag the coverage up to watch the blind spot close — a pass rate over the probed set says nothing about the rest.

6

Reviewed by opendroid · 2026-08-18