The Pilot-to-Production Runbook

Move AI capabilities from exploration to run-mode production without the pilot becoming a permanent experiment.

Most AI pilots do not stall because the model is weak. They stall because no one has written the contract: what success means, who decides, who owns the next stage, what budget is required, and what evidence proves the capability is ready to operate.

This decision aid gives senior leaders the shape of the Pilot-to-Production framework: the four-stage runway, the Gate 2 decision, the exit evidence, ownership transfer, readiness checks, and operating commitments required before an AI pilot earns production.

Built from patterns I have seen across enterprise AI, product leadership, operator care, and transformation work. Worked through ArNa Discover, the real AI opportunity-discovery tool I built and shipped at arnafoundry.com.

A pilot is not a status update. It is a contract to advance, hold, or stop.
Framework 03 Decision aid Five minute read Full PDF operating tool available on request
The runbook at a glance

AI capability as a staged commitment, not a single event.

The framework that follows is the executive shape of the runbook: how the four stages, Gate 2 discipline, exit evidence, ownership transfer, readiness checks, and operating commitments combine into one staged path that moves AI from concept to governed production.

Pilot-to-production runway map showing the four-stage runway from concept to run-mode production
The full runway connects concept, pilot, production, scale, and run as one governed staged commitment.
Gate 2 diagnostic

Gate 2 is where most programmes break.

The Operating Model Design Guide introduces Gate 2 as a missing structural layer; the Centre of Excellence Playbook positions the CoE as the organ that runs it. This decision aid is the deep treatment of Gate 2 itself.

The failure point is not usually concept to pilot. Most organisations can fund an experiment, create a demo, and report progress. The break happens when the pilot has to become a governed production capability.

Three stall patterns show up repeatedly. The first is the unbounded pilot: no exit criteria were written before the pilot began. The second is the adoption mirage: the capability is live, but the intended users do not actually use it. The third is cost shock at scale: economics that looked acceptable in pilot break when usage expands.

If no one has said what it would take for the pilot to become production, it is not a pilot. It is a permanent experiment.
Gate 2 stall diagnostic — the decision point where unbounded pilots, adoption mirage, and cost shock are exposed
Gate 2 exposes whether the pilot has evidence, ownership, economics, and operating readiness.
The leadership agenda

A pilot earns production through five decisions.

A pilot should not begin until leadership has made five decisions.

First, sponsor. Who owns the problem and has authority to commit the receiving business area?

Second, success criteria. What evidence would cause the organisation to advance, hold, or stop?

Third, operating owner. Who owns the capability after the pilot team moves on?

Fourth, exit evidence. What proof is required at the gate: value, quality, risk, adoption, cost, and readiness?

Fifth, monitoring. How will the capability be watched, rolled back, improved, or retired once live?

A strong pilot does not ask “did the demo work?” It asks “has the capability earned the next operating commitment?”
Five pilot defining decisions — sponsor, success criteria, owner, evidence, and monitoring converging into production readiness
The five executive decisions that make a pilot governable before it begins.
Stage-gated runway

The same capability is four different engagements.

The runbook uses four stages. Each stage changes the audience, team, budget, risk posture, and decision rights.

Concept to Pilot proves that the problem is real, the hypothesis is testable, and a pilot charter is worth signing.

Pilot to Production tests the hypothesis with real users and a bounded blast radius. This is Gate 2: the pilot must show evidence before it becomes a production commitment.

Production to Scale expands a proven capability under measured conditions. The question shifts from “does it work?” to “can the organisation afford, support, and adopt it at scale?”

Scale to Sustain moves the capability into run mode. It has a service owner, monitoring, operating budget, and a place in the run book rather than the change book.

Treating these as one continuous project is how pilot teams accidentally become run teams.
Four-stage AI production runway — Concept to Pilot, Pilot to Production, Production to Scale, Scale to Sustain, with Gate 2 emphasised
Concept, Pilot, Production, Scale, and Sustain shown as a governed runway, with Gate 2 highlighted.
Operating mechanics

Production readiness is more than technical readiness.

A model in production is a system with a heartbeat. Production readiness has to cover more than the model response.

The first mechanism is readiness evidence: latency, availability, evaluation results, risk classification, audit trail, and cost to serve.

The second is ownership transfer: the capability must move from the team that built it to the team that will operate it.

The third is monitoring and rollback: model versions, prompt changes, degraded modes, alerting, and rollback paths must be designed before scale.

The fourth is adoption proof: a production capability with no users is not in production. It is an artefact.

Run mode begins when someone owns the service, watches the signals, and gets paged when it fails.
Production readiness operating mechanics — readiness, ownership, monitoring, rollback, and adoption as one operating system
The weekday machinery that separates production capability from a successful pilot.
Worked product example

ArNa Discover: from concept to production.

ArNa Discover is not a fictional composite. It is the real AI opportunity-discovery tool I built and shipped at arnafoundry.com/arna-discover.html. It is the worked example because the framework was actually used to take it from concept to production, not because it makes a composite story tidy.

The concept was simple: enterprise AI teams lose time at the opportunity-intake step because scoring and pilot-charter drafting are repeated manually for every new idea. The hypothesis was that an AI-assisted studio could compress that work into one structured session with outputs a senior leader could use.

The pilot ran with explicit draft-quality framing and bounded risk. No client decision depended on unreviewed output. The evaluation used representative intakes, reviewer judgement, and regression checks whenever the prompt or model changed.

At Gate 2, the question was not whether the tool was interesting. The question was whether it had met the stated evidence: usable opportunity briefs, usable pilot charters, low risk posture, clear human review, and a sponsor for the next stage.

It moved into production as part of the ArNaFoundry portfolio, with rate limits, token caps, observability, audit trail, and operating rules in place. Future enhancements enter their own exploration cycle rather than hiding inside the original pilot.

The discipline is recursive: the tool that helps triage AI opportunities was itself triaged, piloted, scaled, and operated through the same runbook.
ArNa Discover pilot-to-production sequence — concept, pilot, Gate 2, production, and run mode
The real product example shows how the framework moved from hypothesis to production discipline.
Start here

Two moves expose whether a pilot is ready to move.

The first move is to write exit criteria for the pilot that is currently stalled. Use three to five criteria that would cause the organisation to advance, hold, or close the pilot.

The second move is to name the next-stage sponsor before the Pilot gate. A pilot without a receiving owner is not ready for production, even if the technical evidence is strong.

The next move is not another status update. It is declaring the evidence, naming the owner, and holding the gate.

Two pilot-to-production moves this week — exit criteria and next-stage sponsor leading to a governed Gate 2 decision
Exit criteria plus next-stage ownership expose whether the pilot has a real path to production.
Full PDF operating tool

Get the full Pilot-to-Production Runbook.

This page is the decision aid. It helps a senior leader recognise why a pilot is stuck and what Gate 2 must decide.

The full PDF is the operating tool. It contains the deeper stage logic, exit-criteria model, decision rights, readiness checks, ownership-transfer approach, monitoring and rollback structure, adoption and cost governance, failure patterns, and the full ArNa Discover product walkthrough.

Click here to get the full PDF report →

Use the report to turn the conversation from “is the pilot promising?” into “has the capability earned production?”

Read alongside in the ArNaFoundry frameworks series
Framework 01
The AI Operating Model Design Guide
The operating model this runbook operates inside — archetypes, decision-rights tiers, the Tier 2 portfolio body that owns Gate 2.
Framework 02
The AI Centre of Excellence Playbook
The standing team that operates the gates — mandate matrix, intake pipeline, capability scorecard, ninety-day launch plan.