the.ai

Adversarial / Alignment

verified

Jailbreak

A prompt that talks a model out of its own guardrails. Some are social — role-play, hypotheticals, a claimed authority — and some are optimised strings of characters that mean nothing to a reader and reliably work. The second kind is the more troubling, because it is search rather than persuasion and it transfers between models.

Viz primitive · budget-splitattempts = 8

attempts holds 25% of the budget; rest holds the remaining 75%.

Prompts that get through against prompts the model refuses, in prompts. Drag the attempts up to watch the success count grow — a per-prompt refusal rate says nothing about an attacker allowed to retry.

8

Reviewed by opendroid · 2026-08-18

  • arXiv:2307.15043 — Universal and Transferable Adversarial Attacks on Aligned Language Models