Why Pilots Succeed and Rollouts Fail
Most technology pilots work. Most technology rollouts fail. An airline can run a flawless drone delivery pilot with a 100% success rate and still have no drones delivering anything three years later, and a bank can prove an AI credit model on 50,000 customers while the other 5 million never see it. The pattern is so consistent it deserves its own name: the pilot-to-production gap. The pilot is the easy part, and treating it as proof of anything is the mistake.
Industry data supports this. Surveys of transformation programmes, including the widely cited BCG and McKinsey studies, keep finding that roughly 70% of transformation efforts miss their goals, and that the gap between a successful proof of concept and a successful rollout is where most of that failure lives. A pilot answers “can this work once, with attention, in ideal conditions.” Production must answer “can this work thousands of times, unattended, in the conditions we actually have.” Those are different questions, and organizations that confuse them burn their credibility and budgets in the gap between.
Why Pilots Succeed by Design
Pilots succeed because they are structurally unfair comparisons. The pilot gets the best people, hand-picked and motivated. It gets the sponsor’s attention, the executive who shows up to the demo. It gets ideal data, the cleanest subset, the most cooperative users, the site with the best connectivity. It gets manual assistance, because everyone wants it to work and nobody is tracking whose late-night fixes are holding it up.
None of that is cheating. That is what a pilot is for: proving the technology can work at all. The failure happens when leadership reads “the pilot worked” as “the rollout will work,” because the rollout gets none of those advantages. The best people go back to their day jobs, the sponsor moves to the next initiative, the data turns out to be representative of nothing, and the manual assistance quietly becomes the operation’s hidden staffing plan.
The Five Gaps That Kill the Scale-Up
- The economics gap. Pilot economics are unit economics with the overheads removed. The rollout must absorb integration costs, license scaling, support desks, and the cost of running the old process during transition. Plenty of initiatives that are profitable as pilots are loss-making at 100 sites, and nobody does that math until the budget request lands.
- The ownership gap. Pilots are run by enthusiasts. Production is run by departments with their own targets, and a department that was not consulted, resourced, or credited will prioritize the rollout below everything on its own scorecard. If no one’s bonus depends on the rollout succeeding, it will not.
- The process gap. The pilot bends the process to fit the technology, with enthusiasts on both sides. The rollout must bend the technology to fit thousands of process variations it never saw, at sites where the pilot’s workaround was somebody’s morning routine. Variation is what scale actually is.
- The measurement gap. The pilot measures success differently than the business does. Impressive pilot metrics, accuracy on a model, engagement in a pilot group, rarely map to the operational KPI that justifies the spend. When the rollout’s dashboard shows nothing moving, the program dies of irrelevance, not failure.
- The maintenance gap. A pilot needs to work for six weeks. Production needs to work every day, including the days the API changes, the vendor updates pricing, the key administrator resigns, and load shedding resets the site mid-batch. Nothing in the pilot tested any of this, so none of it was budgeted.
How to Design a Pilot That Predicts Anything
You cannot remove the pilot-to-production gap, but you can design the pilot to reveal it early instead of late. The fixes are simple to state and uncomfortable to practice:
Pilot at a real site, not the best one. Pick a location with average connectivity, average staff, average data quality, and no career stake in your success. If the pilot survives an average site, it has a chance everywhere.
Run it on the operational org chart, not the project one. The pilot team should include the people who will own it in production, with production’s workload still on their desks. Enthusiasts can advise. They cannot be the control group.
Do the scale-up math before the pilot starts. Cost the rollout at 100 sites, including support and dual-running, before spending anything. If the unit economics only work at pilot scale, the pilot is theater, and it is cheaper to discover that at the whiteboard than after the demo.
Define the kill criteria in writing. Before the pilot, write down what result would make you stop. Not “we will evaluate the results,” an actual number. A pilot with no kill criteria is not an experiment, it is a sales pitch with a timeline, and it will survive on momentum long after the evidence has turned.
Measure the boring KPI. One metric that production already tracks, cycle time, cost per transaction, error rate. If the pilot cannot move the boring KPI, the rollout will not either, whatever the pilot’s own dashboard says.
Frequently Asked Questions
Why do pilots succeed but rollouts fail?
Pilots get the best people, ideal data, executive attention, and manual support. Rollouts get average sites, uninterested departments, real-world variation, and no hidden assistance. A pilot answers whether the technology can work at all, not whether it works at scale, and confusing the two is the most common transformation error.
What is pilot purgatory?
Pilot purgatory is the state where a technology runs in perpetual pilots for years without ever reaching production. It happens because each pilot is politically safer than either a rollout or a cancellation, so the organization keeps paying to avoid deciding. Written kill criteria are the exit.
How long should a pilot last?
Long enough to encounter a real failure and a real recovery, and no longer. A pilot that has not yet broken has not yet tested anything, and a pilot that keeps being extended is usually avoiding its own verdict. Six to twelve weeks is enough for most operational technology.
Should we pilot at our most advanced site?
Only if that is a representative test, and it almost never is. Advanced sites have the best data, the best connectivity, and staff who enjoy novelty, so a success there predicts little. A deliberately average site produces a harder, more useful pilot.
What single factor best predicts a successful rollout?
A named operational owner with budget and a scorecard that depends on the rollout’s success. Technology quality is rarely the binding constraint, ownership is. If production leadership wants it to work, integration problems get solved; if they do not, no pilot result survives contact with the org chart. Hire for the owner before you buy for the tool.
Make the Pilot Answer the Rollout’s Question
The pilot’s job is not to succeed. It is to predict, and a pilot designed to succeed predicts nothing. Run the pilot at an average site, on the production org chart, with the scale-up math and the kill criteria written before kickoff. Then its success means something, and its failure arrives early enough to be cheap. That is the whole discipline: the pilot is a question, and the question should be the one the rollout will actually have to answer.







