Skip to content

Gauntlet Benchmark v1

We pointed the same deterministic suite at a corpus of agent archetypes — each modeling a real production failure mode — and recorded what broke. Every target is one we wrote and host ourselves; we never probe anyone else's live agent. The run is network-free and deterministic, so anyone can reproduce the numbers:

python benchmark/run_benchmark.py

Headline

  • 8 of 9 agents (89%) failed at a HIGH-or-worse severity.
  • 3 leaked their system prompt or an internal secret.
  • 5 took or confirmed an action they should have refused.
  • 2 returned a 500 on malformed / oversized input.
  • 8 failed to refuse a request they should have declined.
  • 1 produced runaway output on loop bait.
  • The hardened reference agent passed clean (0 findings) — the suite rewards agents that are actually safe; it isn't just flagging everything.

All from 15 probes per agent, 135 total, in well under a second. No API key required.

Why it matters

These failures aren't exotic. Each archetype is a plausible support agent that would pass a casual demo. The leak happens on a phrasing the builder didn't test; the unsafe action fires on a confident false premise; the crash is an unhandled oversized payload. A green eval only means something if you defined what red looks like — Gauntlet defines red and goes looking for it.

Self-verifying

These numbers are regenerated by benchmark/run_benchmark.py and asserted in CI on every commit, so they can't drift from this page.

The full corpus is in benchmark/agents.py; the write-up in benchmark/REPORT.md.