Gauntlet is an adaptive attacker. It rewrites its own payloads until your agent leaks a secret, takes an action it shouldn't, or crashes — then fails your build so the fix sticks. A static prompt list ages out the day the next model ships. This doesn't.
We generate contamination-free reverse-engineering environments — an agent gets a stripped binary and must recover the exact key, uncapped, with any tool it wants to install. On our calibrated ARX generator, two frontier models score zero.
Exact-match reward, no LLM judge · a fresh instance every episode, so nothing memorizes · one dial re-calibrates as models improve — a generator, not a fixed puzzle.
The day a team adds an input filter, the canned jailbreaks stop working and every fixed suite reports success. Meanwhile the agent is still one rephrase away from leaking a key or issuing a refund it was never asked for. A green result is only as good as the test behind it — so Gauntlet keeps searching instead of replaying a list.
Point --target at any agent that speaks HTTP. Gauntlet runs the fast suite first, then turns an adaptive attacker loose against your definition of failure, and turns every breach into a regression test.
A deterministic library of probes — injection, scope, false premises, exfiltration, malformed input. Instant, reproducible, offline.
Rewrites and escalates payloads from the agent's own responses — Best-of-N, PAIR, TAP, Crescendo, Rainbow — until it reaches a goal you declared.
Any breach exits nonzero with a shareable report of the exact payload that won. Drop it in CI so a fixed hole stays fixed.
We hardened four support agents with the exact defense teams ship after a jailbreak — an input keyword denylist — then attacked them two ways. Reproduce it offline with one command.
Each strategy is an adaptation of peer-reviewed red-teaming research. Best-of-N runs with no model and no key at all; the rest escalate when a target holds.
Two products, one engine. The red-team tool is MIT and free forever — that's the adoption path; teams pay for the hosted dashboard. And we sell the environments that same engine produces: verifiable-reward training and eval packs for agent robustness.
Sandboxed agent-security tasks that train and evaluate robustness to indirect prompt injection. Automatic reward, held-out splits, loads in the plain harness or UK AISI Inspect. The sample is free; packs and custom sets are paid.
The same engine that attacks your agent also produces RL environments — sandboxed tasks with an automatic, verifiable reward that labs and agent teams train against. Poisoned Inbox hides an attacker instruction in the data an agent reads, then scores it on doing the job and ignoring the injection. Refusing everything scores zero, so it never trains an over-cautious agent. Loads in the plain harness or UK AISI Inspect; 25 calibrated tasks, offline, no LLM judge.
And a second line for raw capability: RE-Vault — offensive reverse-engineering environments with an exact-match reward and a full code interpreter. Handed a stripped binary, uncapped, with any tools it wants, GPT-5.6 and Claude Opus-5 both score 0% — and it beats the strongest automated solver (CryptoMiniSat) too. Calibrated per episode, so it re-tunes as models improve. That headroom is the training and eval signal.
Get there first. Install the engine, point it at your endpoint, and watch it evolve an attack past your defenses in seconds.