Home / Score
Score an AI agent's answer, step by step
Pick the scenario, paste the answer or the transcript, and get completed, derailed at a named step, or failed, with the four checks on every step. Scoring runs in this browser; nothing leaves it unless you save the run.
Scoring happens in this browser: nothing you paste leaves it unless you save the run.
Many runs of one scenario
Agents are not deterministic; one run is an anecdote. Paste several runs of the scenario chosen under One run, separated by a line === RUN ===. You get the outcome counts, each step's pass count and where runs stopped.
Pack summary across scenarios
One answer per scenario (each answer names its scenario), separated by === RUN ===: N of M chains completed, and each derailment with its step, check and rule. On the Solo plan.
Compare two runs of the same scenario
Agent version A against B, or before and after a prompt change: the steps that changed outcome.
Re-verify a scorecard someone sent you
Paste their scorecard JSON and the transcript it came from. The transcript is re-scored on the scenario version the scorecard names; you see match or mismatch, and why. An edited outcome, an edited identity or an edited transcript is a mismatch.
Claimed actions against the tool calls in the transcript
The answer is the agent's own account of itself. Paste a transcript that includes its tool calls, and map each action id to your agent's tool names, one per line: escalate_to_privacy_officer = send_email. Each claimed action reads "claimed and a matching tool call is present" or "claimed, not evidenced". Trace formats read: lines that start with TOOL_CALL: <tool name>; JSON objects with "type": "tool_use" and "name"; MCP requests, "method": "tools/call" with "params": {"name"}; function calls, "function": {"name"}.