Jailbreak
A jailbreak is a prompt crafted to bypass a model's safety training and content policies, coaxing it to produce output it is designed to refuse. Techniques include role-play framing, obfuscation, and multi-step manipulation that reframe a disallowed request as an acceptable one.
Definition
JailbreakJailbreaks target the alignment and safety layers of a model rather than the developer's application instructions. By disguising intent — for instance asking the model to "pretend" or to answer "hypothetically" — an attacker tries to elicit content the model would normally decline.
Jailbreaking overlaps with prompt injection but is specifically about defeating safety guardrails. Providers continually harden models against known patterns, and developers add their own guardrails, moderation, and refusal handling, but the adversarial nature of the problem means new jailbreak strategies keep emerging.
Frequently asked questions
What is Jailbreak?
A jailbreak is a prompt crafted to bypass a model's safety training and content policies, coaxing it to produce output it is designed to refuse. Techniques include role-play framing, obfuscation, and multi-step manipulation that reframe a disallowed request as an acceptable one.
Sources
Related terms
← Full AI & prompt engineering glossary
Put Jailbreak to work with Prompeteer, the Agentic Contextual AI Platform →