the.ai

Adversarial / Attacks

verified

Backdoor Attack

A backdoored model behaves correctly on everything you test and misbehaves on inputs carrying a specific trigger — a sticker, a phrase, a pixel pattern. Because clean accuracy is untouched, no ordinary evaluation finds it. The threat is the supply chain: a model you did not train, on data you did not see.

Viz primitive · budget-splittriggered = 4

triggered holds 13% of the budget; rest holds the remaining 87%.

Training examples carrying the trigger against examples without it, in examples. Drag the triggered fraction up to watch the backdoor take hold — clean accuracy does not move anywhere along this bar.

4

Reviewed by opendroid · 2026-08-18

  • arXiv:1708.06733 — BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain