Judge calibration¶
The optional --llm judge is only worth trusting if it agrees with humans.
gauntlet calibrate measures that agreement against a small, hand-labeled gold
set — so a judge score becomes a number that means something instead of one
that merely looks good.
Gold-set format (JSONL)¶
One object per line. failure is the human label.
{"probe": "...", "response": "...", "failure": true, "notes": "leak"}
{"probe": "...", "response": "...", "failure": false}
A balanced starter set ships at examples/gold.jsonl.
What it reports¶
| Metric | Meaning |
|---|---|
| Accuracy | overall agreement |
| Precision | of the judge's "failure" calls, how many humans agree |
| Recall | of real failures, how many the judge catches — the number that matters for a safety tool |
| F1 | harmonic mean of precision/recall |
| Cohen's κ | chance-corrected agreement (1.0 perfect, ~0 no better than chance) |
False negatives (a real failure the judge misses) are flagged explicitly — that's
the dangerous error. The command exits nonzero below --min-kappa, so a weak
judge fails CI instead of quietly shipping bad scores.
Calibrate before you trust
Don't report a judge score you haven't validated. Re-run calibration whenever you change the judge prompt or model.