Book cover of Tales from AI Development Hell by Paul Forrest beside the headline Your AI Project Is Fine. Until it isn't.

Your AI Project Is Fine. The Dashboard Says So.

Why ambitious organisations keep building artificial intelligence disasters, and what to do before yours becomes somebody else’s case study.

There is a particular moment in the life of an AI initiative when everything appears to be going splendidly.

The business case has been approved. The vendor has been selected. The steering committee has steered. An ethics assessment has been completed by somebody ethical. The dashboard is green, the launch date is fixed, and an executive has begun using the phrase “transformational capability” without visible discomfort.

Somewhere else in the building, a person who actually understands the system is quietly wondering why nobody has asked what happens when it gets something important wrong.

That gap is the subject of my new book, Tales from AI Development Hell: Moonshots, Misfires and Management — How to Avoid an AI Boondoggle.

It is about what happens when perfectly sensible people, working within perfectly respectable processes, collectively produce an outcome none of them intended and everybody subsequently describes as unforeseeable.

Usually, it was foreseeable. Somebody foresaw it. They simply lacked the seniority, vocabulary or suitably animated PowerPoint slide to interrupt proceedings.

Failure is a management system

A system is bought before its purpose has been defined. Its success is measured using a convenient number that bears only a passing acquaintance with the desired outcome. Warnings become less specific as they travel upwards. Accountability is distributed so efficiently across committees that nobody retains enough of it to do anything useful.

Eventually, somebody outside the organisation experiences the consequences. This is when the organisation discovers that “we followed the process” is not the reassurance it sounded like internally.

The thesis is straightforward: spectacular AI incidents are rarely technology stories. They are the visible end of a quieter management system built from incentives, procurement decisions, misplaced confidence, inadequate measurement and organisational silence.

Moonshots, misfires and the moment things become a boondoggle

Not every failed AI project is a scandal. In fact, an organisation that tests a bold idea, discovers it does not work and stops has achieved something unusually sophisticated: it has learned.

A moonshot is a genuinely ambitious attempt to create disproportionate value. Better cancer treatment, easier property transactions and more accessible public services are all objectives worth pursuing.

A misfire happens when the evidence reveals that reality has failed to cooperate. Models behave differently in production. Historical data encodes historical prejudice. Customers ask questions nobody included in the demonstration. This is not a moral failing. It is useful information, provided someone is prepared to receive it.

A boondoggle begins when the organisation continues to scale, defend or conceal the project after the evidence has changed.

In other words, failure is not the scandal. Failed learning is.

When the dashboard improves as reality deteriorates

Consider Zillow, whose valuation model helped power a business that bought houses directly from homeowners.

As an online estimate, an imperfect house valuation is mostly an invitation to an argument over dinner. As a binding cash offer, the same imperfect valuation becomes an unusually efficient mechanism for buying the homes you have most overvalued.

Sellers decline the offers that are too low. They enthusiastically accept the ones that are too high. The model’s respectable average accuracy does nothing to protect a business that selectively acquires its most expensive mistakes.

Zillow ultimately recorded a $407.9 million full-year inventory valuation adjustment. The lesson is not that predicting house prices is silly. It is that attaching a prediction to a financial commitment changes the entire system, whether or not the slide deck notices.

Elsewhere, Air Canada discovered that a chatbot could provide incorrect bereavement-fare information and that a tribunal was unimpressed by the suggestion that the airline should somehow be less responsible because the answer came from its own digital channel.

The book also examines cancer-care technology, recruitment systems, the Dutch childcare-benefits scandal, facial recognition, fabricated legal authorities, municipal chatbots, autonomous vehicles and software capable of committing money faster than its owners can intervene.

The sectors differ. The organisational choreography does not.

Again and again, the measurement improves while the underlying harm grows. A chatbot’s containment rate rises because distressed customers disappear. Alert volumes look like adoption when they are actually evidence of false positives. Revenue increases while financial exposure becomes steadily more deranged.

The dashboard remains green because nobody asked it to measure the thing catching fire.

The person missing from the meeting

The most important question in the book is also the one most organisations manage not to ask:

What happens to the person the system gets wrong?

Not the person approving the budget. Not the vendor. Not the transformation director whose bonus has acquired an unfortunate dependency on the launch date.

The customer incorrectly advised about arrears. The applicant screened out because yesterday’s hiring decisions became tomorrow’s training data. The family accused by an automated risk score. The pedestrian whom a reassuring safety metric has somehow failed to accommodate.

Can the harm be reversed? Can the affected person challenge the decision? Is there an actual human with authority to overturn it? Does anybody monitor outcomes rather than transactions? Can someone stop the system tonight, without convening a committee whose next available meeting is Thursday?

If the answer is “we have a governance framework”, the question has not been answered.

Nine things that work better than another committee

This is not simply a catalogue of disasters. That would be entertaining, but approximately as useful as a fire-safety manual consisting entirely of photographs of flames.

The book offers a practical operating model: choose, prove, bound, test, own, observe, contest, stop and learn.

Choose problems according to the consequences of getting them wrong, not just the attractiveness of automating them. Prove what the system can actually do before making claims about what it might eventually achieve. Bound its authority, financial exposure and potential harm. Test under realistic conditions, including inconvenient ones.

Give ownership to a named person. Observe what happens after deployment. Make decisions contestable. Establish the authority and mechanism to stop. Then learn from what the evidence tells you, particularly when it tells you something your business case would prefer not to hear.

The accompanying toolkit includes an Evidence Ladder, decision-exposure maps, autonomy budgets, practical controls and a board-assurance checklist. Most fit on one page. This is deliberate. An elaborate control that nobody uses is indistinguishable from no control, except that it has usually cost more.

Five questions for Monday morning

If your organisation is investing in AI, ask five questions about every consequential system it operates:

  1. What does this system decide, and what did we decide it should optimise?
  2. What have we proved, and what are we therefore allowed to claim?
  3. What is the most it can cost us, and who set that number?
  4. Who can stop it tonight, without asking?
  5. What happens to the person we get wrong?

The opposite of AI development hell is not technological perfection. It is competent operations: ambitious, observable, bounded, accountable and capable of changing course.

Which may not sound quite as exciting as a moonshot.

It does, however, tend to be cheaper than the crater.

Buy Tales from AI Development Hell: Moonshots, Misfires and Management — How to Avoid an AI Boondoggle on Amazon.

Read it before your organisation becomes the material for a sequel.