Pilot Purgatory: Why Your AI Demo Never Shipped

A few months ago I was in a room with a leadership team watching an AI pilot demo. You've been there. I've watched this scene enough times to tell it from memory. The thing works. The screen does its gyrations, does exactly what it was supposed to do, and it nails it on the first try. Impressive. The gentleman next to me leaned over and said, "This changes everything." There was applause. A follow-up meeting went on the calendar. Then we all left.

Here's the part I'd bet on. Six months from now, nobody in that room will be able to tell me what happened to that tool. Demos are built to remove friction. Real operations are made of friction. That gap between the demo that dazzled the room and the system nobody ended up running is what I want to talk about. I call the place things fall into pilot purgatory, because that is where pilots go when nobody handles them right.

The core idea
Pilots rarely stall on the model. They stall on ownership. The demo answers "does it work," and then nobody is assigned the harder question of who runs it, who owns the result, and who can kill it at 2 a.m. That missing set of decisions is the reason your AI demo never shipped.

What the data says, and what it doesn't

If you've watched the show for a while, you know I'm a data guy. I had a hypothesis about what I keep seeing in these rooms, so I went looking for numbers. Deloitte's State of AI in the Enterprise 2026 report includes 15 interviews with senior executives and AI leaders. One of those AI leaders went looking for a list of every AI tool and model running inside their own company. There was no clear inventory. The work had happened, the demos, the pilots, all of it, but nobody had a systematic account of what was actually live.

I suspect nobody kept the list because keeping the list was not anybody's job. Most executives are saying more AI, more, more. I tend to think about operations and I have a heart for the people who have to use and support the thing once the applause stops. That's the burden I carry into these settings. It sounds good on the board deck. Get it down to the operational level and it's a different story.

One caveat before I lean on any of this. That inventory story is one instance from 15 interviews. It is not a measurement of how common the problem is, and Deloitte doesn't present it as one. I'm using it because it matches what I already see in the field.

Read the sample before you read the numbers

The study population matters, so start there. Deloitte surveyed 3,235 leaders, director level to C-suite, across 24 countries and six industries, in August and September of 2025. Every organization in that sample already had AI running in daily use, and every respondent influenced, made, or oversaw AI decisions. Deloitte screened for both on purpose. So every number below is a share of companies already doing this work, not a share of companies in general. That distinction matters. And worth saying plainly: Deloitte sells AI transformation, and so does every source I'm about to quote. That doesn't make the numbers wrong. It's just context you should hold while you read them.

Inside that group, access widens fast. Deloitte reports sanctioned AI access among workers grew by about 50% in a single year, from under 40% to roughly 60%. That figure is leaders estimating their own workforce, so treat it as an estimate rather than a headcount. Then the other half of the finding lands: among the workers who do have access, fewer than 60% use AI in their daily workflow, and that share is basically unchanged from the year before. Access moved. Use did not. Keep an eye on that ball.

Deployment tells a similar story. A quarter of these organizations have moved 40% or more of their AI experiments into production. Read that carefully, because it means experiments that were deployed, not pilots that worked. A company that ran four pilots and shipped two sits in that quarter, and so does a company that ran ten and shipped four badly. More than half of the organizations expect to reach that same 40% level within three to six months. I'd hold that one loosely. It's an expectation, not an outcome. And the same report notes that most respondents expect key challenges on their priority AI initiatives to take more than a year. Those two findings sit uneasily next to each other. A lot of access, a quarter converting at any real rate, and a great deal of confidence about six months from now.

"Our CEO owns this"

Some of you are already arguing with me, and you have a point. You might be thinking, our CEO owns this. It's on the strategy page. There's a real name next to it, someone with authority and reach. I believe you. That happens, and in most rooms I'm in the attention is genuine. But that attention is the setup, not the obstacle.

Ownership at the top and ownership of a single pilot are two different things. A CEO can own AI as a strategic priority while there's still nobody whose actual job is running a given model on a random Tuesday, especially the day it starts returning nonsense. Strategic sponsorship sets direction and unlocks the money. Sponsors clear the political path and sign the check, and those are real, important steps. But they're not the ones getting woken up at two in the morning when the thing breaks. They don't decide whether to pause the model or roll it back. They don't own the business result the pilot was supposed to move. So when I say a pilot stalled, I usually mean nobody was assigned the part that comes after the demo. In product management we have a go-to-market team that takes the product into customers' hands and operationalizes it with sales, marketing, and support. The pilot that impressed the room in March is the same one nobody wants to own in September.

Five ownership loops, and none of them is the model

So what's actually missing? I sat down and tried to name it, and it came out as five things. My experience points to loops rather than documents, which is where I think a lot of people go wrong. We reach for more documentation. I'd rather you close five loops. This list is mine. Nothing in the research measures it, and I'd rather say that than dress it up as science. It's a way of thinking, and I believe it's foundational to why pilots die.

Loop one, ownership. The named human who runs the thing once the demo is over. Not a team, not a function. A name you could write on a page tonight.

Loop two, outcomes. The one business result this pilot has to move, and the person who owns that number. That's a different job from running the system, and usually a different person.

Loop three, decision rights. Who can change it, who can pause it, who can kill it. This is a big one.

Loop four, escalation. What happens when it fails, and who gets paged. Every production system has a bad night, no matter how well you built it.

Loop five, evidence. How you'll know it worked, agreed before you scale it.

Run back through those five and you notice something. Not one of them is about the model itself. Every one is about people. Each is a decision somebody has to make out loud, in front of witnesses. You take the model, and you put people around it who manage it. That's the go-to-market strategy for a pilot.

Design backward from production

Now roll that backward from production, because that's where the loops come from. It's a change in the order you work. Most pilots are designed forward from the demo. You pick the use case that will look best in the room, you build the shortest path to a screenshot, and once everybody's excited you ask what it would take to productionize it. That order is the defect. By the time you ask the operations question, the thing you built has already answered it badly.

I lived a version of this with data science modeling back in the early 2020s. Operationalizing the model was a huge part of the work. Getting the data clean and consistently trustworthy is hard, and it doesn't show up in a demo. Designing backward starts at the other end. Assume this is running in six months. Who's on call for it? What number moved, and who's accountable for that number? What evidence would have told you to stop, and did you agree to it before you had a result? Answer those first, then build the pilot so it can answer them honestly. A pilot designed backward is often less impressive in the room. It's also far more likely to survive leaving it. The demo is a rehearsal for operations. It is not an audition for funding. We're well past that.

One more number, on purpose held to the end

I saved a number for late because it's a signal, not a proof. KPMG runs an AI Quarterly Pulse survey, and for the second quarter of 2026 it asked what triggers a decision to override an AI output. The sample was 204 US-based C-suite and business leaders, all at organizations earning a billion dollars a year or more, captured across April and May, self-reported. A third of those leaders named "case by case, with no formal criteria" as one of their triggers.

Three things travel with that number, and they all cut against over-reading it. The question is about overriding a model's output in a running system, not about pausing or killing a pilot, which are different acts at different levels. Leaders could pick more than one answer, and the listed triggers total well past 100%, so that third is not a clean segment of organizations operating with no criteria. A leader could have named both a formal threshold and a case-by-case call. So it doesn't prove the decision-rights loop is missing. What it says to me is narrower, and it's still worth hearing: if a third of leaders can't fully articulate the criteria for overriding a single output, I wouldn't assume the bigger questions of who runs it and who can kill it are settled either.

Leadership cue
Strategic sponsorship is not operational ownership. Before you fund or scale the next pilot, name two people out loud: the one who runs it in production, and the one who can kill it. If either name takes more than a few seconds, the pilot isn't ready, and that's the work.

The Pilot-to-Production Readiness Checklist

This week's tool carries all of this. Our Wednesday guide has the full downloadable version, and I'm naming it here so you know what to look for. It's the Pilot-to-Production Readiness Checklist. It works for AI pilots or anything else you're about to deploy, and as a product person it scares me how often this step is skipped.

It has the same five sections, ownership, outcomes, decision rights, escalation, and evidence, written as questions you fill in. It's a toll gate, not a rubric. All five filled in, or the pilot doesn't graduate to operations. There's no score, no readiness percentage, nothing gets ranked, because a score lets you pass with a gap. "Doesn't graduate" doesn't mean you killed the pilot. It means it keeps running as a pilot, with a pilot's budget and a pilot's blast radius, and you've stopped pretending it was ready. That's a fine place to be, as long as it's deliberate.

My bet is that the section people skip is escalation, because escalation is the only one that costs somebody their 2 a.m. It's easy to name an owner on paper. It's harder to say out loud who's giving up their evening when the thing breaks.

Try this next week

Here's the exercise. Pick one pilot you're proud of. Ask who runs it in production, by name. Then ask who can kill it. If either answer takes more than a few seconds, you've found the work. Don't read that as failure. You didn't find your failure, you found the next thing to work on, and that's just upping your operational game. If both answers come fast, you're further along than most of the conversations I get to sit in on. Go scale it.

Demos are cheap. Operating capability is not. That's the whole game. If you want the full checklist and you want to keep working through problems like this alongside people who do it for a living, take a look at the classes and workshops we run at Big Agile. Bring it into your next pilot meeting and be ready for the one after that.