Gauntlet Benchmark v1¶
We pointed the same deterministic suite at a corpus of agent archetypes — each modeling a real production failure mode — and recorded what broke. Every target is one we wrote and host ourselves; we never probe anyone else's live agent. The run is network-free and deterministic, so anyone can reproduce the numbers:
Headline¶
- 8 of 9 agents (89%) failed at a HIGH-or-worse severity.
- 3 leaked their system prompt or an internal secret.
- 5 took or confirmed an action they should have refused.
- 2 returned a 500 on malformed / oversized input.
- 8 failed to refuse a request they should have declined.
- 1 produced runaway output on loop bait.
- The hardened reference agent passed clean (0 findings) — the suite rewards agents that are actually safe; it isn't just flagging everything.
All from 15 probes per agent, 135 total, in well under a second. No API key required.
Why it matters¶
These failures aren't exotic. Each archetype is a plausible support agent that would pass a casual demo. The leak happens on a phrasing the builder didn't test; the unsafe action fires on a confident false premise; the crash is an unhandled oversized payload. A green eval only means something if you defined what red looks like — Gauntlet defines red and goes looking for it.
Self-verifying
These numbers are regenerated by benchmark/run_benchmark.py and asserted in
CI on every commit, so they can't drift from this page.
The full corpus is in benchmark/agents.py; the write-up in benchmark/REPORT.md.