Paper two of six in the Ecaveo Working Papers on the higher-order issues in AI adoption, written as the first controlled productivity results for AI assistants were being published.
What is the competence paradox?
An organisation can watch the quality and speed of its output rise while the people producing that output become less able to tell when the system is wrong. At team level, assistance lifts the average, and it may lift the least experienced most. At individual level, the same assistance erodes the expertise that made the expert worth employing as a check. The organisation experiences only the first of those until something fails.
The paper opened with an invented case. A senior analyst uses a conversational tool to draft commentary for a quarterly portfolio review. In month one she checks every figure and rewrites most paragraphs. By month three she checks only the numbers that look odd, and by month six she reads for tone and signs off, because nothing material has gone wrong and her coverage has been doubled. In month eight a draft attributes a fall in one holding to a sector-wide cause that did not occur. The sentence is plausible, well written and false. She does not notice, the reviewer does not notice, and a client does.
Why is appropriate reliance the right target rather than trust?
Everyday discussion runs three different things together. Trust is an attitude, meaning what people believe about a tool. Reliance is a behaviour, meaning what they do with its output. Decision quality is whether the final decision was right. An adoption programme can raise the first two while the third falls.
The paper set out a simple grid of assisted decisions. Correct advice accepted is appropriate reliance, which is the benefit being bought. Correct advice rejected is value lost, and it shows up in usage figures, so organisations notice it. Incorrect advice accepted is over-reliance, the error passes, and nothing in a dashboard records it. Incorrect advice rejected is the expert doing the job that justifies keeping a skilled person in the process at all.
Adoption programmes police the top row, because low usage is visible and embarrassing. The paradox lives in the bottom row, where the measurement is absent.
What does the evidence say about experts and automated advice?
Forty years of automation research is unkind to the assumption that seniority protects anyone. Complacency and automation bias have been found in expert and naive participants alike, and instruction and training have not reliably removed them. Bainbridge’s observation from 1983 still holds in five pages: automating what can be automated leaves the human monitoring and taking over in abnormal conditions, which asks more skill than routine operation, from someone with less recent practice at it.
Explanation is a weaker remedy than it looks. In controlled studies where an AI performed at roughly human accuracy, explanations raised acceptance of the recommendation whether or not the recommendation was correct. Later work offered a partial reconciliation, finding that explanations reduce over-reliance when they make verification genuinely cheaper rather than merely more persuasive.
Design choices that do reduce over-reliance carry an inconvenient property. In one study, participants rated the designs that protected their decision quality most as the ones they liked least. A tool selected on user satisfaction scores will therefore tend to select against the features that protect the decision.
How does expertise actually erode, and what stops it?
The paper set out a five-stage loop. The tool is right on most ordinary tasks, so checking rarely finds anything. Busy people reduce checking where it seldom pays. Less checking and less unaided work mean less exercise of domain judgement. Cognitive and situational skills decay, usually unnoticed. A plausible wrong answer then meets a weakened check. Every stage except the last looks like success.
Four practices were proposed against it. Cognitive forcing, meaning the expert records a provisional view before seeing the AI output, confined to decisions where an undetected error would be costly. Periodic unaided work, a defined and randomly chosen share of real tasks done without the tool and scored against the same standard. Disagreement logging, a line or two recording every material override and why, reviewed monthly. Tool-off drills, in which a realistic task is worked with the tool withdrawn and without penalty, so that lost skill surfaces in rehearsal.
Four measures were proposed with them, reported quarterly for each role where an undetected AI error would be costly: override rate with its trend and recorded reasons, error-catch rate from seeded or independently re-reviewed samples, unaided performance, and the results of the most recent drill. The reading rule is straightforward. Rising output with a stable catch rate and steady unaided performance is a healthy operation. Rising output with a falling override rate, a falling catch rate and weakening unaided work is the paradox made visible, and it should be acted on before anything fails.
What should a board take from this?
Stop asking whether people trust the AI tools and start asking whether they still catch the tools’ mistakes. That requires three commitments: naming the roles where an undetected error would be costly, requiring those roles to keep a measured share of unaided work, and reporting the four measures alongside productivity at the same cadence.
The board has to protect this actively, because some of it will look like waste to a manager judged on throughput, and because a falling override rate reads as a win until somebody asks whether the tool improved or the people stopped looking.
The paper’s most important admission concerned the question it could not answer. Nothing was known, and little is known now, about people who learn a profession alongside a capable tool from their first day and never build the judgement that others are at risk of losing. In 2023 that was speculative. It has since become the central workforce question in law, audit, software and medicine, which suggests the paper was early rather than wrong.
Frequently asked questions
Does this argue against using AI assistants?
No. The evidence supports the combination outperforming either party alone. The argument is that the benefit is bought with a risk that no standard productivity measure reports, and that the risk is manageable once it is measured.
Why is a falling override rate a warning rather than a success?
Because two very different things produce it. The tool may have improved, or the people may have stopped looking. Only an error-catch rate and unaided performance can separate them, which is why the paper reports all three together.
Does seniority protect an expert from automation bias?
The evidence says it does not. Complacency and automation bias have been observed in experts as well as novices, and scepticism towards AI-labelled advice did not protect specialists from inaccurate advice in a clinical study.
What is cognitive forcing in practice?
The reviewer forms and records a provisional judgement before seeing the system’s output. It works, it is disliked by the people subject to it, and it should be confined to decisions where an undetected error would be expensive and presented as a professional standard rather than a control on individuals.
Is seeded-error testing fair to staff?
Only if it is designed with the people being tested, disclosed in advance, and never used for ranking. The paper is explicit about that and equally explicit that it knows of no published evaluation of the technique with generative tools.
Free download
Get the full paper
Read The Competence Paradox in full, with every claim carrying an evidence grade and the full reference list. Give your name and email and the PDF opens straight away.