Playbook · 21 pages

How do you know your AI assistant works?

Our field guide for turning qualitative assessment into decision-grade evidence: test sets grounded in real conversations, deterministic checks, AI judges calibrated before they earn trust, and release gates your own teams are prepared to sign off on. Expert judgment is the scarcest input, so the system is engineered to spend it only where it materially changes the outcome.

The playbook cover

Get the playbook

The PDF downloads straight away, and we send a copy to your inbox.

Inside the playbook
The playbook foreword: what evaluation is for
The five-verdict grading vocabulary, including correct refusal
The review room: sealed judge verdicts and keyboard grading

What you will learn

01

Build the instrument

Read real conversations, turn the failures into a taxonomy, and write test items that grade themselves.

02

Five verdicts, refusal included

A vocabulary that separates a correct refusal from a missed answer, and points each verdict at the layer to fix.

03

Size it by expert time

The arithmetic most projects skip: 25 items across seven runs is 175 expert judgments, or three to five weeks of calendar.

04

Hire the judge

Calibrating a judge against human labels, and the agreement numbers that are realistic.

05

Review without spreadsheets

Routing, sealed machine verdicts, and ten seconds of mechanics per judgment.

06

Freeze, gate, sign

Versioned sets, thresholds with provenance, OWASP-mapped security items, and approvals with expiry dates.