What the Evidence Must Show Before Launch

Building and Testing Dependable AI Solutions · 5 min read

An AI workflow does not become dependable because it produced good answers in a demonstration. It becomes dependable when someone has tried to break it in the ways that matter, recorded what happened, and compared the result with a standard agreed beforehand. This lesson sets out what leaders should insist on seeing before a workflow touches real customers, tenants, patients or money.

Routine cases show little

Most test sets begin with typical inputs because they are easy to collect. A tool that handles typical inputs well has shown that it is not obviously broken, and nothing more. The costly failures sit elsewhere: the unusual request, the input with a missing detail, the case where a wrong answer causes real harm. A useful test set therefore adds three kinds of case to the routine ones. Unusual cases are legitimate inputs that look different from the norm, such as a scanned form, a mixed-language message or two documents that contradict each other. Missing-information cases leave out something the workflow needs, to see whether the tool says so or guesses. High-consequence cases are those where an error is expensive, however rare they are in practice.

Each of these should be tied to a written requirement. If a requirement says the tool must flag missing data rather than guess, the test set needs cases that exercise exactly that. Without them the requirement is untested, and a good overall score says nothing about it.

Volume does not repair this. Thousands more routine cases mostly repeat what is already known, and a large overall figure can hide a poor result on the small group of cases that matter. Ask for results by type of case, not only in total.

Criteria are set before the test

Pass criteria should be written down, agreed with the accountable owner and dated before the first run. A criterion names a metric, a threshold, and what happens if the threshold is missed. Setting the bar afterwards invites a quiet drift toward whatever the tool happened to achieve. The EU AI Act points the same way for high-risk systems: Article 9(6)-(7) refers to testing against metrics and probabilistic thresholds defined in advance. Outside that scope the principle is simply good practice, and it costs very little.

What is not a test

Three things are often offered in place of testing. The first is expert impression. When two experienced colleagues read some outputs and say they sound credible, they have judged fluency and tone. Fluent text can be wrong, and AI output is fluent by default. Expert time is better spent designing hard cases and checking each output against the source file with an agreed rubric.

The second is more of the same input, covered above. The third is production: treating corrections from live users as the remaining tests. For cheap, reversible errors, learning from use can be reasonable. For costly errors it moves the testing onto the people who are harmed, and users often fail to notice a polished mistake. Human review is a control that sits on top of testing. It is not a replacement for it, because reviewers who see fluent output tend to over-trust it.

Sample size and honest reporting

A pass rate is only as good as the sample behind it. A dozen cases chosen by the builder cannot support a percentage. A sample drawn from one office does not describe a whole network. Report counts alongside rates, say how the cases were chosen, and state which groups of cases the result does and does not cover. The sentence 'passed four of four' is honest. A claim of one hundred percent accuracy built on it is not.

Worked example

Fenwick Housing Association, an invented landlord, builds a tool that sorts incoming tenant repair requests into urgent and routine and drafts a first reply. The team tests it on a few hundred typical requests and reports that nearly all are sorted correctly. The sponsor asks three questions. Which requirements did this test cover? The requirements say that any mention of a gas smell, no heating for a vulnerable resident, or water near electrics must be urgent, and that a request without an address must be returned for clarification. None of those cases were in the set. What was the bar? Nobody wrote one before the run. Who chose the cases? The builder did.

The sponsor does not reject the tool. She asks for a set of urgent and ambiguous requests written with the repairs supervisors, a group of requests with no address, a threshold for urgent cases set in writing before the run and agreed with the head of repairs, and a scoring pass by someone outside the build team. The first tidy result becomes the starting point of the evidence instead of its conclusion.

As of 24 September 2026.

Sign in to save your progress

You can read every lesson without an account. Signing in keeps your place and unlocks the assessment.