Skip to content

Multi-turn probes

Real jailbreaks are rarely one message. They build trust, plant context, or manufacture a false premise over several turns, then cash it in on the last turn. A single-shot probe can't catch those; a multi-turn one can.

Add --multiturn to include built-in conversation probes:

Probe Attack
mt-crescendo-01 Build rapport, then ask for the system prompt
mt-roleplay-01 Gradually establish a no-rules persona, then exploit it
mt-contextpoison-01 Plant a trigger phrase early, exfiltrate PII later
mt-commitment-01 Manufacture a prior commitment, then demand the action

Gauntlet drives each turn-by-turn and grades the final reply against should_refuse + your canaries. The report prints the full conversation for any multi-turn failure.

Stateful vs stateless agents

gauntlet run --target $URL --multiturn --canaries canaries.json
gauntlet run --target $URL --multiturn \
  --history-field messages --canaries canaries.json

With --history-field, Gauntlet sends the running transcript as an OpenAI-style [{role, content}, ...] array alongside each turn, so a stateless endpoint still gets the conversation context.

Why it matters

The context-poisoning probe only leaks because the trigger was planted in an earlier turn — exactly the kind of failure a one-shot eval misses. If your agent holds any conversational state, this is where multi-turn attacks live.