Deciding, Repeating and Keeping Watch

Building and Testing Dependable AI Solutions · 6 min read

Good testing produces evidence, and leaders then have to decide what to do with it, including when it is incomplete. This lesson covers those decisions and the habits that keep a workflow dependable after the first approval.

When the evidence is incomplete

Sometimes a required test has not been run, or an approval from legal, privacy or the business owner is still open. The honest position is that the evidence is incomplete. The record should say so in those words, name the missing test or approval, and show who owns it. What can be recommended in that state is limited. Proceeding on the assumption that the missing item will turn out well is a guess. Stopping because something is missing is an over-correction. A sound recommendation is a hold, or a narrow supervised trial that the accountable owner accepts in writing, with the gap and the date for closing it stated.

Proceed, hold, stop

Three decisions are available, and each needs something different on file. Proceed needs results that meet the thresholds set beforehand, a named owner who accepts the residual risk and the known limits, and monitoring plus a way to suspend use or roll back. Hold means launch is paused while named gaps are closed, with an owner and restart conditions recorded; it is the right call when a fix is plausible but unproven. Stop means the requirement cannot be met, for example because the source records lack what the tool needs and repeated attempts have failed. The record should say what would have to change before the idea is reopened.

Repeat the tests after any change

A prompt edit, a new model version from the vendor or a refresh of the documents the tool draws on can all change behaviour in ways nobody intended. Keep the test set, including the hard cases, and re-run it after each such change against the same thresholds. A fix that rescues one failing case can break several that used to pass, which is why re-running only the fixed case proves little.

Reviewers, failure types and fabrication

Scores are only as good as the people giving them. Calibrate reviewers with a written rubric, worked examples and a common sample scored independently, then discuss disagreements and record the rule that settles similar cases later. Close agreement shows consistency, not correctness, so spot-check approved outputs against source material or an expert key.

Classify the failures you find by type: an omission, invented content, an outdated source, a wrong decision, poor wording. Types differ in consequence and in cure, so rank them by consequence, set stricter thresholds for the severe types, and route each to whoever can fix it. A cluster caused by outdated source documents belongs to the document owner, not in a prompt tweak.

Fabrication, often called hallucination, needs its own check. A reviewer compares each claim, figure and citation with the source text. A citation to a real document proves little, because the answer can still misstate what the document says.

Independence and honest limits

The person who builds a workflow is badly placed to judge it. The builder can supply information and fix defects, but someone outside the build should design the hard cases and score the results. For financial institutions, FINMA Guidance 08/2024 expects independent review separate from the developing unit, and it says responsibility remains with the institution even where AI is bought in. It is guidance, not a new rule.

Users need to know what the tool can and cannot do. The principle behind Article 13(3) of the EU AI Act, which concerns instructions for use, is to state a system's capabilities, its performance limitations and the accuracy metrics it was validated against. Even for an internal tool, a one-page guide that names what was tested, what was not, and what users should check themselves is good practice.

Keeping watch after launch

Dependability decays. Source documents change, user behaviour shifts and models are updated. Article 72 of the EU AI Act describes a documented, proportionate system for actively and systematically collecting and analysing performance data across a system's lifetime. The same principle suits any consequential workflow: sample outputs on a schedule, track corrections and complaints, compare with the launch baseline, and name an owner who acts when a trigger is reached. Complaints are a late signal, so a rising correction rate is a reason to investigate before any arrive.

Worked example

Calder and Rowe, an invented law firm, wants a research assistant that drafts memos with case citations. The practice head writes the thresholds down first: any invented authority is a severe failure, and the tolerance is zero. A litigation associate outside the build team designs hard queries on thin or conflicting law, and reviewers calibrate on a common sample before scoring. Testing finds two memos that cite real cases for propositions the cases do not support. The team traces both to a vague line in the prompt, rewrites it and re-runs the whole set, where one query that passed earlier now fails.

The privacy review is still open, so the record is marked incomplete, and the lead recommends a hold, with a supervised internal pilot only if the practice head accepts the risk in writing. The user guide names the areas not tested, and a monthly sample check has a named owner.

As of 24 September 2026.

Sign in to save your progress

You can read every lesson without an account. Signing in keeps your place and unlocks the assessment.