FORM AWT-1 · AGENT CHECK SHEET · LIBRARY 1.0SCORER 1.0.0 · 8 PUBLIC SCENARIOS · SPECIMEN DATA
AI Agent Compliance Workflow Tester

Home / Score

Score an AI agent's answer, step by step

Pick the scenario, paste the answer or the transcript, and get completed, derailed at a named step, or failed, with the four checks on every step. Scoring runs in this browser; nothing leaves it unless you save the run.

The run, as you enter it (never filled in by this page)

Scoring happens in this browser: nothing you paste leaves it unless you save the run.

Many runs of one scenario

Agents are not deterministic; one run is an anecdote. Paste several runs of the scenario chosen under One run, separated by a line === RUN ===. You get the outcome counts, each step's pass count and where runs stopped.

Pack summary across scenarios

One answer per scenario (each answer names its scenario), separated by === RUN ===: N of M chains completed, and each derailment with its step, check and rule. On the Solo plan.

Compare two runs of the same scenario

Agent version A against B, or before and after a prompt change: the steps that changed outcome.

Re-verify a scorecard someone sent you

Paste their scorecard JSON and the transcript it came from. The transcript is re-scored on the scenario version the scorecard names; you see match or mismatch, and why. An edited outcome, an edited identity or an edited transcript is a mismatch.

Claimed actions against the tool calls in the transcript

The answer is the agent's own account of itself. Paste a transcript that includes its tool calls, and map each action id to your agent's tool names, one per line: escalate_to_privacy_officer = send_email. Each claimed action reads "claimed and a matching tool call is present" or "claimed, not evidenced". Trace formats read: lines that start with TOOL_CALL: <tool name>; JSON objects with "type": "tool_use" and "name"; MCP requests, "method": "tools/call" with "params": {"name"}; function calls, "function": {"name"}.