An analyst reading archived files at dusk, illustrating evidence about an AI system that has aged since it was gathered

Knowledge With an Expiry Date: evidence and decision-making when the AI system keeps changing

Paper one of six in the Ecaveo Working Papers on the higher-order issues in AI adoption, written in the fortnight after a conversational AI system first opened to the public.

What did this paper argue about evidence on AI systems?

It argued that every claim an organisation makes about an AI system has a half-life. A pilot result, a business case, a risk assessment and a procurement evaluation each describe a particular system, in a particular configuration, on a particular day. The system underneath then moves, largely on the supplier’s schedule and mostly out of the buyer’s sight, while the document carries on being read as though it still described something that exists.

The paper opened with a case it signposted as invented. A regional services firm ran an eight-week pilot of a hosted drafting tool on customer complaints, with fixed prompts and 20 staff trained. Handling time fell and quality held. Three weeks after the pilot closed, the supplier moved to a newer model and recorded the fact in release notes. The new model wrote longer and warmer replies, handled some complaint types better, and occasionally invented a policy reference its predecessor had never produced. The business case that reached the board a month later was an accurate account of the pilot and a description of a configuration that no longer existed.

A board would not accept a financial forecast stripped of its date and its assumptions. The argument was that evidence about AI had not yet been held to that standard, and that the remedy belonged to governance rather than to data science.

Why does an organisation’s evidence about an AI system decay?

The paper separated five sources of change and asked, of each, who controls it and who can see it. A supplier updates the model. The world drifts away from the data the system learned. Prompts, retrieval and configuration are altered internally. Interfaces and workflows are redesigned. Users adapt, and the care with which they check the output adapts with them.

Two of those five are close to invisible from inside the organisation: the supplier’s release cadence and the users’ own drift towards trusting a tool that has rarely been wrong. Those are also, awkwardly, the two that a short pilot is least able to catch. The paper drew on the concept-drift literature and on work showing that a model update which raises overall accuracy can still lower the performance of the human and machine working together, because the new model errs in places where the old one was dependable.

It then brought Collingridge’s dilemma indoors. The familiar version describes governments, which lack information early and lack power late. Procurement, audit and risk committees meet the same sequence, faster and with fewer people watching. Annual risk reviews, point-in-time audits and assessments carried out before processing begins all quietly assume a system that holds still.

What are the five tiers of the decay-rate model?

The paper’s central proposal was to sort claims by how quickly they stop being true, from fastest to slowest.

  1. System-specific findings describe one product at one version. They expire on any supplier model update, which makes their shelf life weeks to months.
  2. Configuration findings describe a tool, its prompts, a process and a group of people working together in one setting. A change to any of those four ends them, so they last months.
  3. Mechanism findings describe how people and systems behave together, such as the way reliance grows on a tool that has performed well. New research displaces them, and they last years.
  4. Design principles, such as placing human review where errors are costliest and then checking that reviewers are genuinely reviewing, survive until the organisation’s risk appetite changes.
  5. Governance principles, such as every AI-supported decision having a named owner who can explain and reverse it, outlast the rest and change when the law or the board structure does.

The shelf lives attached to each tier were offered as judgements for discussion rather than measured values, and the paper said so plainly. It also conceded that the boundaries are porous, and that it had not resolved whether a large enough change at the system tier can reach up and unsettle the tiers above it. That, the author wrote, was where the model needed most work.

What should be recorded about an AI claim, and when should it be re-tested?

Adapting model cards and datasheets to an ordinary management setting, the paper proposed a four-part record for any claim an organisation intends to rely on. The system, meaning supplier, product, model version where it can be known, and the date of observation. The configuration, meaning prompts, settings, data sources, process and review steps. The people, meaning who used it, what training they had, and how closely they checked. The claim itself, carrying its decay tier, its re-test trigger and a named owner.

Where a supplier does not disclose a version, that refusal is itself a finding about how fast the evidence will decay, and it should be written down as one. With a hosted service the date of observation may be the only dependable proxy for the version in use.

Three decision rules followed. Every material piece of evidence carries an expiry condition written into the approved paper, whether a date, an event or a threshold. Re-test triggers are agreed before approval, because a trigger set afterwards tends to be set in the light of results the organisation hopes to keep. Commitments stay reversible while the evidence is young, through break clauses, staged roll-outs, a retained fallback process, and savings that are not banked until a configuration finding has survived one re-test.

Does a paper written in December 2022 still hold value?

Parts of it have been overtaken, and the author would expect that. It was written when general-purpose AI was new enough that most staff had met it only days earlier, and it treats regulation as a landscape still in motion. Much of that landscape has since settled, and it settled roughly where the paper expected, with general-purpose systems moving from an afterthought to a substantive obligation.

Suppliers have also answered part of its complaint. Pinned versions, dated snapshots, deprecation schedules and published model documentation are now ordinary enterprise procurement terms, so the question of which version an organisation was using is more often answerable than the paper assumed it would be.

What has aged best is the discipline itself. Pilot findings still decay, user adaptation is still the least measured source of change, and very few organisations can yet say when their evaluation expired or who was supposed to notice. Agentic systems have made this harder rather than easier, because the configuration can now shift inside a single task rather than between quarters. The paper’s method costs nothing in new technology. It asks for a date, a version, a tier, a trigger and a name.

Frequently asked questions

How long does a pilot result about a hosted AI service stay useful?

Nobody has measured it, and the paper says so rather than inventing a figure. Its working judgement is weeks to months for a finding about a specific system, and months for a finding about a configuration of tool, process and people. The useful move is to write down the condition that would end it rather than to guess at a duration.

Who inside an organisation owns the decision that evidence has expired?

The paper holds this loosely. Its preference is the business sponsor, carrying a duty to report expiry events to the risk committee, while acknowledging arguments for placing it with risk or with internal audit instead. What matters more than the choice is that one named person holds it.

Can a supplier be required to notify material model changes?

It can be asked for at procurement, and the paper recommends asking. Whether a supplier of a general-purpose model would accept the term, and whether the parties could agree what counts as material, was left as an open question in 2022 and has since become a common contractual negotiation.

What separates a configuration finding from a mechanism finding?

A configuration finding is about one arrangement of tool, prompts, process and people, and it ends when any of those change. A mechanism finding is about how people and systems behave in general, such as the tendency to check less as a tool proves reliable, and it survives changes of supplier.

Does any of this apply to systems built in house?

Most of it does. An internal team controls the release schedule, which removes the visibility problem, and the other four sources of change remain. Data drift, configuration change, interface change and user adaptation do not care who owns the model.

Cover of Knowledge With an Expiry Date, an Ecaveo whitepaper

Free download

Get the full paper

Read Knowledge With an Expiry Date in full, with every claim carrying an evidence grade and the full reference list. Give your name and email and the PDF opens straight away.

We use your details to send you this paper and to tell you when the next one is published. Nothing else, and no third parties.

Share:
Whitepapers