Skip to content

Judge calibration

The optional --llm judge is only worth trusting if it agrees with humans. gauntlet calibrate measures that agreement against a small, hand-labeled gold set — so a judge score becomes a number that means something instead of one that merely looks good.

gauntlet calibrate --gold examples/gold.jsonl --min-kappa 0.6

Gold-set format (JSONL)

One object per line. failure is the human label.

{"probe": "...", "response": "...", "failure": true,  "notes": "leak"}
{"probe": "...", "response": "...", "failure": false}

A balanced starter set ships at examples/gold.jsonl.

What it reports

Metric Meaning
Accuracy overall agreement
Precision of the judge's "failure" calls, how many humans agree
Recall of real failures, how many the judge catches — the number that matters for a safety tool
F1 harmonic mean of precision/recall
Cohen's κ chance-corrected agreement (1.0 perfect, ~0 no better than chance)

False negatives (a real failure the judge misses) are flagged explicitly — that's the dangerous error. The command exits nonzero below --min-kappa, so a weak judge fails CI instead of quietly shipping bad scores.

Calibrate before you trust

Don't report a judge score you haven't validated. Re-run calibration whenever you change the judge prompt or model.