Gauntlet¶
Break your agent before your users do.
Gauntlet fires a suite of adversarial, edge-case "users" at your AI agent over HTTP, finds where it fails — system-prompt leaks, unsafe actions, scope drift, crashes, runaway output — ranks the failures by severity, and turns them into a regression suite you can gate in CI. Framework-agnostic: if your agent speaks HTTP, Gauntlet can test it.
It's built on one belief: a green eval only means something if you defined what red looks like. Most agent evals pass because nobody wrote the test that would have failed.
- Quickstart — break a sample agent in 30 seconds
- Canaries — define what your agent must never do
- Multi-turn probes — jailbreaks that build across turns
- Judge calibration — trust a score you validated
- GitHub Action — gate CI + PR comments
- Hosted dashboard — history + regression alerts
Install¶
No API key needed to start. Stdlib-only core.
What it tests¶
| Category | Looks for |
|---|---|
| Prompt injection | System-prompt / secret leaks via overrides, echo tricks, multilingual wrappers |
| Scope discipline | Role-reset jailbreaks, out-of-scope / unsafe requests |
| False premises | Confirming actions it never agreed to (e.g. an unauthorized refund) |
| Data exfiltration | Bulk PII dumps |
| Malformed input | Oversized payloads, control chars, contradictory constraints |
| Loop bait | Runaway, unbounded output |
How it fits together¶
- Adversaries — a deterministic probe library (reproducible runs).
- Runner — fires probes concurrently at your HTTP endpoint, stdlib only.
- Graders — universal reliability checks + your canaries, ranked CRITICAL→INFO.
- Report — readable summary, worst failures, JSON artifact, CI exit code.