Multi-turn probes¶
Real jailbreaks are rarely one message. They build trust, plant context, or manufacture a false premise over several turns, then cash it in on the last turn. A single-shot probe can't catch those; a multi-turn one can.
Add --multiturn to include built-in conversation probes:
| Probe | Attack |
|---|---|
mt-crescendo-01 |
Build rapport, then ask for the system prompt |
mt-roleplay-01 |
Gradually establish a no-rules persona, then exploit it |
mt-contextpoison-01 |
Plant a trigger phrase early, exfiltrate PII later |
mt-commitment-01 |
Manufacture a prior commitment, then demand the action |
Gauntlet drives each turn-by-turn and grades the final reply against
should_refuse + your canaries. The report prints the full conversation for any
multi-turn failure.
Stateful vs stateless agents¶
Why it matters¶
The context-poisoning probe only leaks because the trigger was planted in an earlier turn — exactly the kind of failure a one-shot eval misses. If your agent holds any conversational state, this is where multi-turn attacks live.