Most AI pilots do not stall because the model is weak. They stall because no one has written the contract: what success means, who decides, who owns the next stage, what budget is required, and what evidence proves the capability is ready to operate.
This decision aid gives senior leaders the shape of the Pilot-to-Production framework: the four-stage runway, the Gate 2 decision, the exit evidence, ownership transfer, readiness checks, and operating commitments required before an AI pilot earns production.
Built from patterns I have seen across enterprise AI, product leadership, operator care, and transformation work. Worked through ArNa Discover, the real AI opportunity-discovery tool I built and shipped at arnafoundry.com.
The framework that follows is the executive shape of the runbook: how the four stages, Gate 2 discipline, exit evidence, ownership transfer, readiness checks, and operating commitments combine into one staged path that moves AI from concept to governed production.
The Operating Model Design Guide introduces Gate 2 as a missing structural layer; the Centre of Excellence Playbook positions the CoE as the organ that runs it. This decision aid is the deep treatment of Gate 2 itself.
The failure point is not usually concept to pilot. Most organisations can fund an experiment, create a demo, and report progress. The break happens when the pilot has to become a governed production capability.
Three stall patterns show up repeatedly. The first is the unbounded pilot: no exit criteria were written before the pilot began. The second is the adoption mirage: the capability is live, but the intended users do not actually use it. The third is cost shock at scale: economics that looked acceptable in pilot break when usage expands.
A pilot should not begin until leadership has made five decisions.
First, sponsor. Who owns the problem and has authority to commit the receiving business area?
Second, success criteria. What evidence would cause the organisation to advance, hold, or stop?
Third, operating owner. Who owns the capability after the pilot team moves on?
Fourth, exit evidence. What proof is required at the gate: value, quality, risk, adoption, cost, and readiness?
Fifth, monitoring. How will the capability be watched, rolled back, improved, or retired once live?
The runbook uses four stages. Each stage changes the audience, team, budget, risk posture, and decision rights.
Concept to Pilot proves that the problem is real, the hypothesis is testable, and a pilot charter is worth signing.
Pilot to Production tests the hypothesis with real users and a bounded blast radius. This is Gate 2: the pilot must show evidence before it becomes a production commitment.
Production to Scale expands a proven capability under measured conditions. The question shifts from “does it work?” to “can the organisation afford, support, and adopt it at scale?”
Scale to Sustain moves the capability into run mode. It has a service owner, monitoring, operating budget, and a place in the run book rather than the change book.
A model in production is a system with a heartbeat. Production readiness has to cover more than the model response.
The first mechanism is readiness evidence: latency, availability, evaluation results, risk classification, audit trail, and cost to serve.
The second is ownership transfer: the capability must move from the team that built it to the team that will operate it.
The third is monitoring and rollback: model versions, prompt changes, degraded modes, alerting, and rollback paths must be designed before scale.
The fourth is adoption proof: a production capability with no users is not in production. It is an artefact.
ArNa Discover is not a fictional composite. It is the real AI opportunity-discovery tool I built and shipped at arnafoundry.com/arna-discover.html. It is the worked example because the framework was actually used to take it from concept to production, not because it makes a composite story tidy.
The concept was simple: enterprise AI teams lose time at the opportunity-intake step because scoring and pilot-charter drafting are repeated manually for every new idea. The hypothesis was that an AI-assisted studio could compress that work into one structured session with outputs a senior leader could use.
The pilot ran with explicit draft-quality framing and bounded risk. No client decision depended on unreviewed output. The evaluation used representative intakes, reviewer judgement, and regression checks whenever the prompt or model changed.
At Gate 2, the question was not whether the tool was interesting. The question was whether it had met the stated evidence: usable opportunity briefs, usable pilot charters, low risk posture, clear human review, and a sponsor for the next stage.
It moved into production as part of the ArNaFoundry portfolio, with rate limits, token caps, observability, audit trail, and operating rules in place. Future enhancements enter their own exploration cycle rather than hiding inside the original pilot.
The first move is to write exit criteria for the pilot that is currently stalled. Use three to five criteria that would cause the organisation to advance, hold, or close the pilot.
The second move is to name the next-stage sponsor before the Pilot gate. A pilot without a receiving owner is not ready for production, even if the technical evidence is strong.
The next move is not another status update. It is declaring the evidence, naming the owner, and holding the gate.
This page is the decision aid. It helps a senior leader recognise why a pilot is stuck and what Gate 2 must decide.
The full PDF is the operating tool. It contains the deeper stage logic, exit-criteria model, decision rights, readiness checks, ownership-transfer approach, monitoring and rollback structure, adoption and cost governance, failure patterns, and the full ArNa Discover product walkthrough.
Use the report to turn the conversation from “is the pilot promising?” into “has the capability earned production?”