Tests, Oversight, Incidents and Change
Governing AI as It Scales · 5 min read
Lesson one settled who owns what. This lesson covers what happens as use grows: how to test before repeating a use, how to tell workers, how to design oversight that works in practice, how to react when something goes wrong, and how to keep control when a vendor changes the product underneath you.
Bounded tests before repeat use
Approve a test, not a rollout. A bounded test has a defined scope, named participants, a permitted list of information that may be used (and an excluded list), an end date, and success measures set in advance. The Act's risk management article asks for testing against metrics and probabilistic thresholds defined in advance (Art. 9(6) and (7)); even where it does not apply to you, fixing the bar before you see results is good practice, because a bar set afterwards drifts to wherever the results landed. Include unusual, incomplete and high-consequence cases, not only routine ones.
A successful test approves that use. A second process, new data types or another team is a new request with its own approval of scope and permitted information. Success in one place is evidence, not permission.
Saying what the evidence shows, and holding when it is incomplete
Leaders are often asked to say a tool is validated. Say only what the evidence supports. If testing covered routine cases and nothing unusual, say that. If privacy approval is still open, say that too. When evidence is incomplete and a required approval is unresolved, the status is hold, and the record is labelled incomplete. It states what was tested, what the results do and do not show, what is unresolved, the conditions to lift the hold, and who decides. Approving conditionally and finishing the privacy review after launch turns an open question into a fait accompli.
Informing workers
If an AI system is used in the workplace, tell the people it affects before it starts. For deployers of systems the EU AI Act covers, Art. 26(7) requires informing workers' representatives and affected workers before workplace use. Where the Act does not reach a use, the step is still sound practice. Tell them what the system does, what data it uses, who owns it, and where to raise concerns.
Oversight that can be exercised
Naming a person as overseer proves nothing. Art. 14(4) describes what an overseer needs to be able to do: understand the system's capacities and limitations and monitor its operation (a); stay aware of automation bias (b); interpret the output correctly (c); decide not to use it, or disregard, override or reverse it (d); and intervene or stop it so it halts in a safe state (e). Translate that into design. Give reviewers the source data, enough time per case, training on how the tool fails, and authority to override without extra forms or manager sign-off. Then test it: seed a few known errors and see whether reviewers catch them. Near-total approval may mean an accurate tool or reviewers who stopped looking; only a check tells you which.
Incident clocks
The Act sets reporting deadlines for serious incidents in Art. 73: 15 days generally, 10 days where a person has died, and 2 days for widespread infringement or serious and irreversible disruption of critical infrastructure. An initial incomplete report followed by a complete one is permitted. The leadership consequence is practical. Decide in advance who judges whether an event is a serious incident, how quickly they can be reached at a weekend, and how information flows between you and your vendor. Escalate first, report what you know, complete the record later.
Vendor changes and the inventory
When a vendor swaps or retrains the model behind a tool you approved, your approval covered the earlier behaviour. Treat the change as a governance event: ask the vendor in writing what changed, re-run your acceptance tests including unusual cases, and have the owner decide before the new version handles consequential work. Responsibility stays with you even where the AI is bought in, a point FINMA Guidance 08/2024 makes for supervised institutions.
All of this depends on keeping an inventory of AI applications, another expectation in that guidance. Each entry names the accountable owner, the risk classification and why, the data used, the approvals held, and the date of last review. An unregistered tool that turns up is added, assessed and decided on, and you ask why it was missing.
A worked example
Fenmoor Regional Hospital Trust is invented. Its appointments team piloted an assistant that drafts patient letters. The test used routine bookings; the privacy officer had not yet approved use of appointment data. The sponsor wanted to announce it as validated. The programme lead instead recorded: routine cases acceptable; unusual bookings untested; privacy approval open; status hold, labelled incomplete. Staff representatives were briefed before any use. A second test then ran on a bounded set of letters, with seeded errors and reviewers empowered to override. When the vendor changed its model mid-test, the lead paused, retested, and the owner decided.
As of 24 September 2026.
You can read every lesson without an account. Signing in keeps your place and unlocks the assessment.