Artificial intelligence in critical national infrastructure has arrived well ahead of the evidence, and every deployment owes an answer to one question, which is what the service does on the morning the model is withdrawn.
Nothing in this article constitutes legal, regulatory or security advice. It is offered to support discussion and further review with qualified advisers.
What is the known safe state for artificial intelligence in critical national infrastructure?
A safe state is where a service goes when the model cannot be trusted. Nothing in the AI Playbook, the Data and AI Ethics Framework, the Responsible AI Toolkit or the transparency standard asks for a fallback to be defined, still less exercised. Write down what the service does without the model, who can run that fallback, how long it lasts, and when it was last exercised.
- Version 4.0 of the Cyber Assessment Framework carries four objectives and 14 principles, none of them specific to artificial intelligence.
- The Algorithmic Transparency Recording Standard binds central government through policy enforced by spend controls, and the official finder returned 152 records across 38 organisations on 12 September 2026, a count that moved between retrievals.
- The cross-government Microsoft 365 Copilot experiment ran with 20,000 participants and no control group, reporting an average of 26 minutes a day saved on a self-reported basis.
- The Court of Appeal held in August 2020 that a police force had breached section 149 of the Equality Act 2010 by never having enquired whether its facial recognition software was biased.
- Five cyber attacks on UK water systems after 1 January 2024 reached the public record only through a freedom of information disclosure, because the NIS regulations trigger reporting on disruption.
How strong is the evidence for AI in government?
Start with the audited baseline. The National Audit Office surveyed 89 government bodies in autumn 2023 with a 98 per cent response rate from 87 organisations, publishing in March 2024, the only such survey by an independent body with a statutory remit. It found 37 per cent of those bodies had deployed artificial intelligence, reporting 74 use cases between them. Of the 32 with something live, eight reported being always or usually compliant with the transparency standard.
Every published evaluation has a design weakness, and the strongest of them reports no saving. The Copilot experiment ran with no control group, and its headline of 26 minutes a day was self-reported, taken from the middle point of each categorical range. Some 17 per cent reported no clear time saving. HM Revenue and Customs allocated licences at random to 3,000 employees, yet the outcome measure remained self-reported time. The most careful study recorded 60 per cent exact-match agreement with human reviewers on consultation analysis, against a human benchmark of 62 per cent measured in internal testing on different data, and gives no cost figure.
Why is public sector AI accountability still a policy commitment?
The Algorithmic Transparency Recording Standard is how government intends to earn public legitimacy for algorithmic decision-making. Its mandatory scope was formalised in a Government Digital Service policy of 17 December 2024. No statute or statutory instrument sits behind it. Enforcement runs through digital and technology spend controls, and the exemptions track the Freedom of Information Act 2000, including commercial sensitivity and the risk of gaming the tool.
Two legal duties sit behind the standard and carry consequences a court can enforce. The Court of Appeal held in R (Bridges) v Chief Constable of South Wales Police that a police force breached section 149 of the Equality Act 2010 by never enquiring whether its facial recognition software was biased. It found no clear evidence of actual bias, and the breach was procedural because the duty is one of enquiry. A vendor’s assurance therefore leaves the duty undischarged, which I would put as reasonably firm, short of certain, since another tribunal on different facts could take a narrower view. Section 80 of the Data (Use and Access) Act 2025 then replaced Article 22 of the UK GDPR from 5 February 2026, relaxing the restriction on solely automated significant decisions while preserving safeguards on information, representations, human intervention and contest. Meaningful human involvement is left undefined, so a department is applying a relaxed rule whose central term it defines for itself.
Where does operational technology security change the deployment question?
Connecting an analytical system to an operational one expands the attack surface. In an office the worst outcome is usually lost or corrupted information, but in an industrial environment it is a physical event. The Secure Connectivity Principles for Operational Technology, published on 14 January 2026 and co-sealed by eight agencies, name real-time analytics, predictive maintenance and remote monitoring as the benefits of connectivity, and record separately that third-party vendors, remote access and supply chain integrations expand that same surface. Putting the two together is my argument and not the guidance’s. The final principle requires an isolation plan that has been tested, which is the safe state in cyber form.
Version 4.0 of the Cyber Assessment Framework, used by competent authorities under the NIS regulations and across the public sector through GovAssure, contains no AI-specific objective, principle or contributing outcome, so an operator can comply fully while running models it cannot explain, on data it cannot trace. The regime covers ten subsectors under Schedule 2 to the Network and Information Systems Regulations 2018, while the National Protective Security Authority lists 13 critical national infrastructure sectors.
Incident visibility follows the same design, since the regime triggers reporting on disruption, so any resilience assessment built on the public record rests on a filtered sample. The Information Commissioner fined South Staffordshire Plc and South Staffordshire Water Plc £963,900 in May 2026 over a breach affecting 633,887 individuals, with roughly 20 months of undetected dwell time and monitoring covering five per cent of the estate. That notice concerns personal data and does not state that operational technology or water supply was affected.
What makes common-mode failure the distinctive public sector risk?
The distinctive risk is many services failing at once because they depend on the same thing. On 19 July 2024 a faulty content update from a single security vendor affected 8.5 million Windows devices worldwide, according to Microsoft, and organisations that had never decided to depend on each other found that they did. The Competition and Markets Authority found in its final cloud report of 31 July 2025 that Amazon Web Services and Microsoft each hold between 30 and 40 per cent of UK infrastructure-as-a-service by value, and in March 2026 it declined to prioritise strategic market status investigations into either firm. It opened one into Microsoft’s business software ecosystem instead, launching in May 2026, a priority it explained partly by the rapid integration of AI into workplace tools.
Where is the value credible?
Correspondence and casework has the strongest fit, and the consultation evaluation is the reason. A measurable ground truth exists, the review step is cheap at a median of 23 seconds per response, and a human benchmark of 62 per cent inter-reviewer agreement supplies a comparator, though it came from internal testing on different data. The Local Government and Social Care Ombudsman handled 22,010 complaints and enquiries in 2024-25, upholding 83 per cent of 4,441 completed investigations.
Asset health sits below casework in the same ordering. A systematic review of railway predictive maintenance, published in March 2026, included 73 studies and found the literature dominated by laboratory and simulation work. It identifies a safety paradox, since railways replace components before failure and produce little run-to-failure data for training. The North Hyde review makes the point. An elevated moisture reading was detected in oil samples taken in July 2018, nearly seven years before the transformer fire that closed Heathrow for most of 21 March 2025, yet mitigating actions appropriate to its severity were left unimplemented. The review does not draw the next conclusion, and I do. Better detection returns little in an organisation that already ignores the detection it has.
What does a safe state and fallback record contain?
One page per service carries it. What does this service do if the model is withdrawn this morning? Who is competent to run it that way, and how many of them are there? For how long can that be sustained, and when was it last exercised with real staff and real volumes? A service that cannot answer all four is arguably unready to deploy, which is my position and stricter than any published guidance.
The exercise is the part that matters, because a written fallback nobody has run is a hypothesis. Take an invented water undertaking that withdraws its pump telemetry anomaly model for six hours on a Wednesday. Two of the four control room staff on shift joined after deployment and have never worked without it, and the manual procedure references a screen layout replaced in 2023. An exercise surfaces that in an afternoon, and an incident surfaces it mid-incident.
Continuity is the constraint these services cannot trade away. A water undertaking cannot suspend supply and a benefits agency cannot suspend payment, so every design decision has to answer what happens when the new component fails. Artificial intelligence in critical national infrastructure is being bought against evaluations that measure self-reported time, under a transparency regime resting on policy, inside an assurance framework written before the current wave.
What follows for an accounting officer is modest. Write down the safe state before buying the model, publish the transparency record early, and instrument the equality duty at deployment. Operational technology security deserves the same treatment, since the NIS regulations surface only what disrupts, and public sector AI accountability becomes real where somebody senior reads an exercise report they found uncomfortable. The safe state and fallback record survives every vendor and every model change.
Frequently asked questions
Is the Algorithmic Transparency Recording Standard a legal requirement?
Cross-government policy makes it mandatory for central government departments and arm’s length bodies dealing with the public, enforced administratively through digital and technology spend controls. No statute or statutory instrument sits behind it. The official finder held 152 unfiltered records across 38 organisations on 12 September 2026, with one sorted query returning 138, so the count is worth re-checking.
Does the Cyber Assessment Framework cover artificial intelligence?
Version 4.0, released on 18 April 2024 and reviewed on 6 August 2025, carries four objectives and 14 principles covering security risk management, protection, detection and impact minimisation. None of them is specific to artificial intelligence.
What did the Court of Appeal decide about algorithmic bias in 2020?
In R (Bridges) v Chief Constable of South Wales Police the Court held that a police force had breached section 149 of the Equality Act 2010, because it had never sought to satisfy itself that its facial recognition software did not have an unacceptable bias on the grounds of race or sex. The breach lay in the failure to enquire.
How much time does artificial intelligence actually save in government?
The cross-government Microsoft 365 Copilot experiment reported an average of 26 minutes a day across 20,000 participants, self-reported and without a control group, with 17 per cent reporting no clear time saving. A randomised licence trial at HM Revenue and Customs measured a self-reported saving of two to three per cent of a working week. No published UK evaluation measures an objective outcome against a control.
What should an accounting officer do first?
Write down the safe state before buying the model. For every service where artificial intelligence will influence a decision or an operation, record what it does without the model, who can run it that way, how long that can be sustained, and when it was last exercised.
Free download
Get the full paper
Read The Known Safe State in full, with every claim carrying an evidence grade and the full reference list. Give your name and email and the PDF opens straight away.