The Pilot-to-Production Checklist
The gap between a successful pilot and a production system isn't scale. It's seven things that the pilot never had to have.
Every AI story follows the same arc. The pilot works — impressively, sometimes miraculously. The demo is great. The stakeholders are excited. And then the pilot ends, and the thing that worked in the pilot has to work in production, where the data is messier, the volume is higher, the users are less patient, and the mistakes cost real money.
Most pilots don't survive the transition. Not because the AI stopped working, but because the pilot was never asked to carry the things production needs. The pilot proved the model could do the task. It never proved the system around it existed.
The seven gates
Gate one: ownership. The pilot had a champion — someone who cared, pushed, and kept it alive. Production needs an owner with a budget and a mandate. Not a sponsor. An owner. The difference is who's accountable when it breaks. If you can't name the person whose job depends on the system working, the system is a pilot forever.
Gate two: evaluation. The pilot was judged by demos and enthusiasm. Production needs a measurement — a test set, a baseline, a defined quality bar that a new model version has to clear before it ships. Without it, every change is a gamble and every regression is a surprise.
Gate three: budget. The pilot ran on free credits, spare time, and goodwill. Production needs a real budget line — for the model calls, the infrastructure, the review labor, and the upkeep. The free pilot is a lie about the cost of the service, and the first bill is the moment the lie gets discovered.
Gate four: an SLA that means something. Production needs a defined service level — accuracy on the things that matter, not just uptime, with a monitoring path and a response plan when it's missed. The pilot had no promises. Production is a promise with a price.
Gate five: support. The pilot had the person who built it. Production has users — people who will have problems, questions, and edge cases at 4pm on a Friday. Production needs a support path: who answers, how fast, and what the escalation looks like. Nobody plans the 4pm Friday until the 4pm Friday happens.
Gate six: training. The pilot was used by the people who built it, who understood its failure modes. Production is used by people who didn't build it, who will trust it too much and too little in equal measure. Production needs training — the operators need to know what it's for, what it gets wrong, and when to override.
Gate seven: rollback. The pilot could be turned off with no consequences. Production can't. It needs a rollback plan — what happens when the new model degrades, when the vendor changes something, when the quality bar fails. The rollback isn't an admission of failure. It's the thing that makes the forward move safe.
A pilot proves the model can do the task. Production is the system around the model. The seven gates are the difference, and skipping any one of them is how pilots become incidents.
Why the gates get skipped
The skip is almost always the same story. The pilot succeeded, the momentum was real, and the pressure to move was overwhelming. Nobody wants to be the person who slows down a successful pilot. So the gates get deferred — "we'll add the monitoring after launch," "we'll document the support path next sprint" — and deferred gates are the ones that never happen, because the incident that forces them is the event that should have had them in place.
The other reason is that the gates are invisible. Ownership, evaluation, budget, SLA, support, training, rollback — none of them show up in a demo. The pilot looks complete because the visible part, the model working, is complete. The system around it isn't visible, so it doesn't get built, until it fails visibly.
How to run the gates without killing momentum
The gates don't have to be a process gauntlet. They're a checklist — ten minutes with the pilot's owner, asking seven questions, and being honest about the answers.
The discipline is the same as any launch review. The gate isn't "is the model good" — that's been proven. The gate is "does the system around the model exist." If the answer to ownership is a name and a budget, pass. If the answer is a shrug, the pilot stays a pilot until it's answered. Not because the process demands it. Because the incident will demand it anyway, and it's cheaper to meet the gate than to meet the incident.
The honest truth about gates: they're not bureaucracy. They're a list of the ways a successful pilot becomes a failed production system. The pilot that clears all seven is rare, and it's rare because most organizations skip the boring gates and discover them the expensive way.
The question that starts the review
When the pilot is declared a success, ask one question: what's different about production?
The answer should be a list — more volume, messier data, real users, real consequences, no champion in the room. Every item on that list maps to a gate. Volume maps to budget and routing. Messy data maps to evaluation and monitoring. Real users map to support and training. Consequences map to the SLA and rollback. The differences aren't abstract. They're the checklist.
Think this argument fits your event? Tell me about the room — the calendar is selective.
Start a conversation