Evaluate an AI assistant with verifiable test cases
A convincing answer is not necessarily correct or usable. Evaluate an assistant on a defined task with retained cases and criteria written before reviewing results.
Define the task and its boundaries
Specify inputs, authorised sources and the result users can use. Separate summarising, recommending, drafting and executing. Define when the assistant should seek clarification or hand over.
Fictional example: an assistant classifies support requests but changes no accounts. Success is a justified category or a clarification request, not a real intervention.
Build a separate case set
Prepare ordinary, ambiguous, incomplete and out-of-scope cases. Add documents containing unwanted instructions to check that the assistant does not treat them as authorisation. Use only fictional or authorised data.
Keep evaluation cases separate from examples used to tune the system. Otherwise apparent improvement may reflect adaptation to familiar examples. Record situation families missing from the set.
Write criteria before testing
For each case, specify required information, forbidden actions and evidence of success. Separate factual correctness, citation quality and scope control. A general style score does not replace these criteria.
Classify errors by their consequences in context: simple correction, poor decision or forbidden action. Agree thresholds with the service owner; they are not a universal standard.
Compare versions under consistent conditions
Record model, instructions, sources, available settings and test date. After a change, rerun the same set and examine differences by case family rather than only an average.
Where results vary, document variation through limited repeat trials. Do not retain only the most favourable answer. Keep representative failures with planned corrections.
Decide using evidence and unknowns
Record authorised use, human checks, excluded cases and shutdown conditions. A small successful set supports at most a bounded trial under tested conditions, not general reliability.
The cited NIST framework provides voluntary risk-management guidance. This guide’s method remains an editorial synthesis, without expert validation or system certification.
A record to keep with the decision
| Item | Evidence |
|---|---|
| Task | Usable outcome, limits and forbidden actions. |
| Case set | Examples, covered families and unknowns. |
| Criteria | Expected outcomes and evidence for each case. |
| Version and decision | Configuration, errors and usage conditions. |
Download the worksheet to fill in (CSV)
Frequently asked questions
How many cases are needed?
The number depends on situations and risks to cover. No quota in this guide establishes general reliability.
Is a good average sufficient?
No. Examine potentially serious failures and underrepresented families separately.
Reference material
NIST · AI Risk Management Framework
The practical checklist is an editorial synthesis to adapt to your service. It does not constitute a certification or an audit result.