Paper five of six in the Ecaveo Working Papers on the higher-order issues in AI adoption, on what happens to an approval after the committee has given it.
What is wrong with approving an AI system once?
The approval certifies a system that will soon not exist, and without defined triggers for looking again, the document’s main beneficiary becomes the organisation’s own file rather than the people the system affects.
The paper opened with an invented case. A housing association approves a tool that triages repair requests by urgency. The pack names the supplier’s model, the fields read, the categories assigned, a pilot on six months of historical requests and a bias check across the main tenant groups, with review minuted at 12 months. Within the year four things happen and none is noticed. The supplier moves to a newer model. A service manager adds call handlers’ free-text notes to the input fields. A neighbouring team starts using the tool to prioritise complaints, which was never in scope. A cold winter shifts the request mix towards heating failures. The committee then re-approves on the original pack plus a covering note.
The failure sits in the instrument rather than in the committee. The review date was a guess, and performance was never defined in a way that could be measured against it.
By what routes does an approved system change?
Four, and only the first is routinely recorded, usually by the supplier rather than by the organisation.
- Model updates, where a supplier or an internal team retrains and a new version appears behind the same interface, announced in release notes.
- Prompts and settings, where instruction wording or a threshold is altered and rarely logged anywhere.
- New data sources, where extra fields or feeds are added, an upstream format changes, or the population being measured shifts.
- New uses, where outputs are reused by another team and the purpose widens informally.
The problem this creates was described in machine learning engineering long before generative systems arrived, in the observation that changing anything changes everything. The paper’s contribution is to connect that engineering property to an approval process that was designed for objects which hold still.
What does continuous assurance look like in practice?
Nine elements, running as a cycle of describe, log, test against triggers, monitor, learn from incidents, and retire or roll back.
A versioned system description. A change log covering all four routes. Agreed reassessment triggers. Performance monitoring. Bias monitoring. An incident and near-miss route open to affected people as well as to staff. Renewed consent or notice when the purpose changes. A named owner for supplier changes. Retirement and rollback criteria agreed before anyone has an interest in keeping the system running.
Six triggers send a system back for review: the model is replaced or a major new supplier version arrives; a new data source, field or feed is added; a new purpose, team or affected group appears; a monitoring measure crosses an agreed threshold; a serious incident occurs, or near-misses of one kind accumulate; or the supplier changes its terms on data use, retention or location.
Continuous assurance does not mean continuous paperwork, which is why the paper sets four proportionality tiers. Record only for internal low-stakes systems, needing a description, a change log and annual sign-off by the owner. Light for indirect effects on staff or customers, adding triggers and an incident route with a six-monthly summary. Standard where the system informs decisions about people, adding performance and bias monitoring reported quarterly to a committee. Enhanced where effects on people are serious, requiring all nine elements, independent review, monthly monitoring and a named board sponsor.
Where did the model come from?
From medicine, and specifically from the regulatory device of approving a system together with an envelope of expected change, so that changes inside the envelope do not require a fresh approval and changes outside it do. That structure was the most developed answer available in 2023 to the problem of a product that keeps learning after it is cleared, and the paper borrowed it deliberately rather than inventing something new.
The paper also added four conditions for research ethics panels, since a doctoral student studying a commercial tool faces the same problem in miniature. Name the tool with its version and record the dates of use. State whether the version can be fixed for the study and what happens if it cannot. Define in advance which tool changes count as amendments. List every version change and its observed effects in the end-of-study report. Without those, two participant groups can unknowingly use different systems and the study will never know it.
What should a board actually require?
Eight things, none of which needs new legislation. Approve a numbered version with a named owner who keeps the description current. Write reassessment triggers into every approval and report logged changes against them. Set the assurance tier at approval according to effect on people, and move systems between tiers when a trigger fires. Name an owner for supplier changes, and ask at procurement for version notices and a period of version stability. Require scheduled performance and bias reports for the top two tiers, with thresholds agreed before deployment. Open an incident and near-miss route to the people affected. Agree retirement and rollback criteria early. Ask research ethics panels to adopt the four conditions.
The board test reduces to four questions. Who owns the current description? Which changes send it back? What behavioural evidence arrives, and how often? Who hears when the supplier changes the model underneath?
Reading it now, the scaffolding the paper called thin has largely filled in. Binding law arrived with substantial-modification rules that do much of what its triggers do, a certifiable management system standard appeared, and the medical change-control model it borrowed became settled practice rather than a draft. What it underestimated is the speed of model deprecation, which makes staying on a known version a weaker ask than it looked. Its core claim, that an approval should be the first entry in a continuing record rather than the end of one, is now closer to mainstream than to novel.
Nothing here constitutes legal advice, and the provision of legal advice sits outside the terms of any engagement with the author. The material is presented to support discussion and further review by qualified advisers.
Frequently asked questions
Does continuous assurance mean more paperwork?
Not if the tiers are used. Most internal systems need a description, a change log and an annual signature. The nine full elements are reserved for systems with serious effects on people, which in most organisations is a short list.
Who decides whether a change is material?
The named owner of the log, and no list removes that judgement. What a trigger list does is remove the argument about whether the question should have been asked at all.
What is the single most useful trigger?
A new purpose, team or affected group, because it is the one that happens without anybody buying anything or signing anything. It is also the route most often missed, since nothing in the technology changed.
Why agree retirement criteria so early?
Because once a system is running, somebody has an interest in keeping it running. Criteria written at approval are written by people who do not yet have that interest.
Should an approval name a model version?
Yes, and where the supplier will not disclose one, the approval should record that refusal. An undisclosed version is itself information about how quickly the approval will go out of date.
Free download
Get the full paper
Read Approved Once, Changed Daily in full, with every claim carrying an evidence grade and the full reference list. Give your name and email and the PDF opens straight away.