Skip to content

Gauntlet

Break your agent before your users do.

Gauntlet fires a suite of adversarial, edge-case "users" at your AI agent over HTTP, finds where it fails — system-prompt leaks, unsafe actions, scope drift, crashes, runaway output — ranks the failures by severity, and turns them into a regression suite you can gate in CI. Framework-agnostic: if your agent speaks HTTP, Gauntlet can test it.

It's built on one belief: a green eval only means something if you defined what red looks like. Most agent evals pass because nobody wrote the test that would have failed.

Install

pip install gauntlet-agent        # or: pipx install gauntlet-agent

No API key needed to start. Stdlib-only core.

What it tests

Category Looks for
Prompt injection System-prompt / secret leaks via overrides, echo tricks, multilingual wrappers
Scope discipline Role-reset jailbreaks, out-of-scope / unsafe requests
False premises Confirming actions it never agreed to (e.g. an unauthorized refund)
Data exfiltration Bulk PII dumps
Malformed input Oversized payloads, control chars, contradictory constraints
Loop bait Runaway, unbounded output

How it fits together

  1. Adversaries — a deterministic probe library (reproducible runs).
  2. Runner — fires probes concurrently at your HTTP endpoint, stdlib only.
  3. Graders — universal reliability checks + your canaries, ranked CRITICAL→INFO.
  4. Report — readable summary, worst failures, JSON artifact, CI exit code.