Skip to main content
arrow_back Back to Hub
Capability

Why AI Pilots Die in the Gap Between the Demo and Tuesday Morning

Weekly dispatch

Stay sharp on AI.

The weekly dispatch for founders, directors and CTOs building real AI capability: the hard-won lessons, not the hype. One email a week.

One useful note each week. Unsubscribe anytime.

The demo that worked

Three months ago I sat in on a demo that went perfectly. An operations director stood up in front of her leadership team and showed an AI agent reading a batch of supplier invoices, matching them against purchase orders, and drafting the query email for the three that didn't reconcile. Four minutes, start to finish. The room was properly impressed. So was I.

By the time I checked back in, the tool had been quietly abandoned. Nobody had switched it off, and nobody had complained about it. It had simply stopped being the thing people reached for on an ordinary Tuesday, and gone back to being the thing that only worked in a demo.

That's not a story about a bad model, or a team that lacked appetite. The model did exactly what it was built to do, in the room where it was shown. Which is precisely why it's worth taking seriously: something can work perfectly in a demo and still be dead within a quarter, and if you don't understand why, you'll keep buying the next demo too, and the one after that.

What the demo doesn't show you

A demo works for a specific, almost engineered reason: every source of friction a real working day contains has been quietly removed before anyone walked into the room.

The data is clean, because someone spent a week preparing a sample set precisely so the model wouldn't choke on a missing field or a supplier name spelled four different ways across four systems. There is exactly one person driving it: the champion who commissioned the tool, understands its quirks, and has learned what to type and in what order to get the good outcome. And there is no competing system in sight. No fifteen other logins. No parallel spreadsheet that someone in finance still trusts more than the new dashboard. No colleague three desks over who's done this task their own way for six years and isn't about to stop now.

None of that is dishonest. It's simply how you demonstrate anything. But it means the demo is a test of the model, not a test of the job. The job is a different animal entirely: a dozen systems that don't talk to each other, data that's clean in one place and a swamp in another, and habits that took years to form and won't unform because a new tool arrived with a good four-minute story.

The pilot doesn't fail when it meets a harder version of the problem it was built for. It fails when it meets a completely different problem, one that was never in the room to begin with.

The work around the work

Here's what actually fills most people's Tuesday: not the invoice, not the report, not the task itself, but everything that surrounds it, purely because the systems and people around it aren't naturally in sync.

It's chasing someone in procurement for a status update because the system doesn't surface it on its own. It's re-keying the same customer detail into three different platforms because nobody ever built, or paid for, the connection between them. It's the meeting that exists only to establish what happened in the meeting before it. It's the Friday afternoon lost to reconciling a spreadsheet that only needs reconciling because two systems refuse to agree on a single number.

None of this appears on an org chart. Nobody's job title is "keeper of the status updates." But add it up across an ordinary week and, for most people in most businesses, it is the majority of the working day: not the visible task, but the invisible labour of keeping everyone and everything roughly pointed the same way. Call it the work around the work.

It stays invisible mainly because it's load-bearing. Stop doing it and things visibly break, so everyone keeps doing it and nobody stops to ask why it's still needed. That's exactly why so many AI pilots aim at the wrong target. They speed up the visible task, the report, the invoice, the summary, while the work around the work, the actual bottleneck, carries on consuming the day exactly as before. You've handed someone a faster way through the twenty per cent of their job that was never the problem, and left the other eighty per cent untouched. No wonder the pilot doesn't feel like it changed anything.

This is not a new ambition

None of this is a new insight, which is worth sitting with. Reconnecting how work actually happens to what the organisation is trying to achieve is one of the oldest ambitions in management. Business process re-engineering tried it in the nineties. Company intranets and knowledge bases tried it in the two thousands. RACI charts, status dashboards, the entire discipline of "knowledge management" all tried some version of the same thing: build a layer that keeps everyone's picture of the work current and shared. We've sat close enough to several of these efforts, on both sides of the table, to have watched the ending arrive in exactly the same shape each time.

They all failed for the same reason. Keeping that layer current was itself a job. Someone had to update the intranet. Someone had to chase the status. Someone had to keep the dashboard honest. That maintenance was a clerical burden layered on top of the work it was meant to illuminate, and it was always the first thing to slip once the initial enthusiasm faded and whoever championed it got busy, changed roles, or left. The tool didn't die because it was a bad idea. It died because nobody could sustain the admin of keeping it alive.

Most AI pilots walk into the same trap wearing new clothes. A pilot that bolts a clever capability onto a process, without removing any of the work around that process, hasn't lightened anyone's day. It's added one more system to check, one more login to remember, one more thing someone has to keep updated for the good outcome from the demo to keep happening. That's an addition to the burden, not a subtraction from it, and it's why adoption quietly dies within a few months even when the model behind the pilot was genuinely good. Nobody sat down and decided to stop using it. Its maintenance simply lost the daily competition for attention, the same competition every well-intentioned dashboard and intranet has been losing for thirty years.

What has to be different this time

If that's the actual pattern, the fix isn't a better model, a more charismatic champion, or a longer training programme. It's designing for the whole day a pilot has to survive, not the four minutes it has to look good in.

That means holding any new capability to three questions before it goes anywhere near a leadership demo, the same three we now hold our own work to before it goes anywhere near a client.

Does it remove more friction than it adds? Not "does it do something impressive" but does it take a real piece of the work around the work off someone's plate, the chasing, the re-keying, the reconciling, rather than simply doing the visible task a little faster while leaving the surrounding admin exactly where it was.

Does it live inside the systems people already use, rather than asking them to adopt a parallel one? A tool with its own login and its own place in someone's attention is competing against years of muscle memory, and muscle memory usually wins.

And, the question every previous version of this idea skipped: who, or what, keeps it maintained once the initial enthusiasm wears off? Every earlier attempt at this depended on a human to keep the connective tissue current: updating the intranet, chasing the status, reconciling the spreadsheet. Humans get busy, get bored, or get promoted into a different job. That isn't a character flaw. It's simply what happens to any task that isn't formally anyone's job.

This is the one genuinely new ingredient AI agents bring to an old problem. Used properly, an agent can be the thing that keeps the connection between systems alive: pulling the status instead of waiting for someone to chase it, reconciling the two spreadsheets on its own, surfacing the update rather than waiting to be asked for one, without needing a human to keep doing that clerical maintenance forever. That's a genuinely different proposition to an intranet nobody updates. It doesn't get bored. It doesn't get busy. It doesn't get promoted into a different job next spring.

It is not, however, automatic, and it is not quick. None of this happens because you switched an agent on. It happens because someone sat down, mapped the actual work around the work in a specific process, and deliberately designed the agent to remove it, which is a slower and far less glamorous job than building an impressive demo. And to be clear what this is for: the point isn't fewer people doing the chasing. It's the same people, freed from the chasing, doing the parts of the job that were the actual reason you hired them in the first place.

Before you run the next pilot

So before the next pilot goes in front of a leadership team, it's worth asking a duller question than "what can this model do?"

What does the work around the work actually look like in the process you're trying to fix? Who's chasing whom? What gets re-keyed, and where? What meeting exists purely to report a status that a system should already know? And, honestly: is the pilot you're about to run designed to remove that layer, or is it just going to sit on top of it as one more thing someone has to remember to manage?

The second version will demo beautifully. It just won't be there by November.

Share