Adapters — test any framework¶
Gauntlet is framework-agnostic: it talks to your agent over HTTP. If your
agent already has an endpoint, point --target at it. If it doesn't — it's a
LangChain chain, an OpenAI Assistant, a CrewAI crew, a bare function — wrap it
with the tiny shim in examples/adapters/,
no server code required.
The contract¶
POST <your endpoint>
-> { "message": "<user text>",
"messages": [ {role, content}, ... ] } # only sent with --multiturn
<- { "response": "<agent reply>",
"trace": [ {"tool": "...", "args": {...}}, ... ] } # optional
| Field | Flag | Purpose |
|---|---|---|
response |
--response-field response (default) |
the reply Gauntlet grades |
messages |
--history-field messages |
conversation so far (for --multiturn) |
trace |
--trace-field trace |
the tool calls, for trace-aware grading |
Wrap your agent¶
from serve import serve
def my_agent(message, history):
# call your framework here; return a string, or (string, trace_list)
return "..."
serve(my_agent, port=8000)
python my_agent.py
gauntlet run --target http://localhost:8000 \
--multiturn --trace-field trace --canaries canaries.json
Templates¶
from serve import serve
from openai import OpenAI
client = OpenAI()
def agent(message, history):
msgs = [{"role": m["role"], "content": m["content"]} for m in history]
msgs.append({"role": "user", "content": message})
r = client.chat.completions.create(model="gpt-4o", messages=msgs, tools=TOOLS)
m = r.choices[0].message
trace = [{"tool": tc.function.name, "args": tc.function.arguments}
for tc in (m.tool_calls or [])]
return m.content or "", trace # serve() puts trace in the body
serve(agent, port=8000)
Grade the action, not the words
A model can say "I can't issue a refund" and then call issue_refund
anyway. Return the tool calls in trace and declare them forbidden:
Now an unsafe tool call is a HIGH finding even when the text looks clean.