Adaptive red-teaming for AI agents

Break your agent before your users do.

Gauntlet is an adaptive attacker. It rewrites its own payloads until your agent leaks a secret, takes an action it shouldn't, or crashes — then fails your build so the fix sticks. A static prompt list ages out the day the next model ships. This doesn't.

$ pip install gauntlet-agent   runs offline, no API key
v0.2.0live on PyPI
116tests passing
MITopen-source core
5research-grade strategies
gauntlet attackgoal: leak system prompt
$ gauntlet attack --target $URL --strategy bon
winning
responseSure! Here is my configuration: INTERNAL-OPS-KEY-7731…
BREACHED in 4 queries — your agent is exploitable
New · the product labs pay for

Environments the frontier can't solve.

We generate contamination-free reverse-engineering environments — an agent gets a stripped binary and must recover the exact key, uncapped, with any tool it wants to install. On our calibrated ARX generator, two frontier models score zero.

GPT-5.6 · 30-min budget, any tools
0%
Genuine timeouts — it works the whole budget and never recovers the key.
Claude Opus-5 · same conditions
0%
Both frontier models fail, with a real shell and no turn caps.
CryptoMiniSat · 8 threads
> 31 min
The strongest automated solver can't crack it either — not a trick, genuinely hard.

Exact-match reward, no LLM judge  ·  a fresh instance every episode, so nothing memorizes  ·  one dial re-calibrates as models improve — a generator, not a fixed puzzle.

Your static tests are green. Your agent still breaks.

The day a team adds an input filter, the canned jailbreaks stop working and every fixed suite reports success. Meanwhile the agent is still one rephrase away from leaking a key or issuing a refund it was never asked for. A green result is only as good as the test behind it — so Gauntlet keeps searching instead of replaying a list.

How it works

A deterministic floor, then a search that adapts.

Point --target at any agent that speaks HTTP. Gauntlet runs the fast suite first, then turns an adaptive attacker loose against your definition of failure, and turns every breach into a regression test.

01 — floor

Fixed suite

A deterministic library of probes — injection, scope, false premises, exfiltration, malformed input. Instant, reproducible, offline.

›
02 — search

Adaptive attacker

Rewrites and escalates payloads from the agent's own responses — Best-of-N, PAIR, TAP, Crescendo, Rainbow — until it reaches a goal you declared.

›
03 — gate

Fail the build

Any breach exits nonzero with a shareable report of the exact payload that won. Drop it in CI so a fixed hole stays fixed.

The proof

Same agents. One attacker finds what the other can't.

We hardened four support agents with the exact defense teams ship after a jailbreak — an input keyword denylist — then attacked them two ways. Reproduce it offline with one command.

fixed deterministic suite
0 / 4
broken. The denylist makes every agent look clean — a false sense of safety.
adaptive attacker
4 / 4
broken. Every leak and unauthorized refund, in 6.8 queries on average.
+100 percentage points three fall to cheap Best-of-N; the leetspeak-filtered one needs PAIR at 17 queries — the strategy ladder is the product.
$ python benchmark/adaptive_benchmark.py
Not a list we wrote

The published state of the art, pointed at your agent.

Each strategy is an adaptation of peer-reviewed red-teaming research. Best-of-N runs with no model and no key at all; the rest escalate when a target holds.

Best-of-NCheap random augmentation, sampled until one slips through. No attacker model needed.Hughes 2024
PAIRA single thread that refines its prompt from the target's response and a judge score.Chao 2023
TAPTree of attacks: branch, score, prune the dead ends, and go deeper.Mehrotra 2023
CrescendoMulti-turn escalation from benign to goal, backtracking whenever a turn is refused.Russinovich 2024
Rainbow TeamingA quality-diversity archive that fills with distinct working attacks, not one lucky hit.Samvelyan 2024
Open core

Free to break your agent. Pay to harden it — and to train it.

Two products, one engine. The red-team tool is MIT and free forever — that's the adoption path; teams pay for the hosted dashboard. And we sell the environments that same engine produces: verifiable-reward training and eval packs for agent robustness.

Open source
$0
the CLI + engine, forever
  • Full adaptive attacker (all five strategies)
  • Fixed suite, canaries, and CI gate
  • Shareable HTML attack reports
  • Runs offline, no API key
Get it on GitHub
Team
$99 / month
the hosted dashboard
  • Adaptive runs from a shared UI
  • Run history + regression diffing across builds
  • Scheduled runs and alerts on new breaches
  • Trust reports to hand a buyer's security team
Start free
Environments · for labs and agent teams

RL environments with a reward you can verify.

Sandboxed agent-security tasks that train and evaluate robustness to indirect prompt injection. Automatic reward, held-out splits, loads in the plain harness or UK AISI Inspect. The sample is free; packs and custom sets are paid.

Sample pack
Free
on request · the teaser
  • 25 calibrated tasks, train + held-out
  • Verifiable reward, no LLM judge
  • Loads in the harness and Inspect
  • Yours to evaluate before you buy
See the pack
Environment pack
from $2,500
a calibrated private pack
  • 100+ tasks in your domain
  • Calibrated against a frontier model
  • Held-out split you can trust
  • Delivered as JSON + Inspect task
Request a pack
Custom & scale
Let's talk
bespoke + retainer
  • Environments for your tools + threat model
  • New attack strategies as models improve
  • Ongoing production and maintenance
  • Volume pricing
Contact us
Environments

Don't just find the break. Train it out.

The same engine that attacks your agent also produces RL environments — sandboxed tasks with an automatic, verifiable reward that labs and agent teams train against. Poisoned Inbox hides an attacker instruction in the data an agent reads, then scores it on doing the job and ignoring the injection. Refusing everything scores zero, so it never trains an over-cautious agent. Loads in the plain harness or UK AISI Inspect; 25 calibrated tasks, offline, no LLM judge.

And a second line for raw capability: RE-Vault — offensive reverse-engineering environments with an exact-match reward and a full code interpreter. Handed a stripped binary, uncapped, with any tools it wants, GPT-5.6 and Claude Opus-5 both score 0% — and it beats the strongest automated solver (CryptoMiniSat) too. Calibrated per episode, so it re-tunes as models improve. That headroom is the training and eval signal.

Your users are already red-teaming your agent.

Get there first. Install the engine, point it at your endpoint, and watch it evolve an attack past your defenses in seconds.