The demo took six weeks. The model summarized documents, answered questions, drafted responses that sounded like a competent analyst wrote them. The steering committee was impressed. Someone said the word transformational. Budget for the next phase was approved in the same meeting.
That was eighteen months ago. Nothing is in production.
This is the most common failure pattern in enterprise AI right now, and it is remarkably consistent across industries. The pilot works. The pilot impresses. Then the pilot dies somewhere between the demo environment and the production environment, slowly enough that nobody has to announce its death. It just stops appearing in the updates.
The instinct is to blame the technology. The model was not accurate enough, the data was not ready, the vendor oversold. Sometimes that is true. Usually it is not. Most AI pilots that die in the gap between demo and production were never going to cross it, because the conditions for crossing it were never built. The pilot was designed to prove the technology. Nobody designed the path.
The demo is not the product
A demo runs on curated data. Someone selected the documents, cleaned the inputs, chose the questions in advance. The audience is friendly and the stakes are zero. When the model produces something odd, the presenter moves to the next example and nobody remembers it by the end of the meeting.
Production is a different object entirely. Production means real data with real edge cases arriving in real time. It means error handling, monitoring, rollback procedures, a support model, and an answer to the question of what happens the day the system is confidently wrong in front of a customer, a patient, or a regulator.
None of that is a model problem. All of it is a delivery problem. And almost none of it exists at the moment the pilot gets its applause, because none of it is needed to produce the applause.
A demo proves the technology works. Production requires proving the organization works.
The distance between those two proofs is where the eighteen months go.
Nobody owns the path
Look at who built the pilot. An innovation team, a data science group, a vendor with a proof-of-concept budget. Now look at who has to run the thing in production. IT operations, security, compliance, the business unit whose workflow it touches, the data platform team whose pipelines feed it.
Every one of those groups can veto the path to production. None of them owns it.
This is the structural gap that kills most pilots, and it is invisible during the pilot phase because the pilot does not need any of those groups to succeed. The pilot has a team. Production needs an owner. Those are different things, and the difference only becomes visible when the pilot ends and the question of who takes it forward has no answer.
The test is simple. Ask who owns this system when it breaks at two in the morning a year from now. If the answer is a discussion instead of a name, the pilot does not have a path to production. It has a demo environment and a hope.
"It looks good" is not an evaluation
Pilots get approved on impressions. The outputs looked right. The summaries read well. The stakeholders who tried it liked it.
Production requires criteria. What error rate is acceptable, on which categories of input, measured how, by whom, reviewed at what cadence. What does an unacceptable output look like, who is accountable when one reaches a user, and what is the correction path.
Generative systems raise this bar instead of lowering it. The failure mode of a traditional system is visible: it crashes, it returns nothing, it throws an error. The failure mode of a language model is a fluent, confident, well-structured answer that is wrong. Evaluation has to be designed to catch exactly the failures that are hardest for a human reviewer to notice, and that design work is real work. It does not happen as a side effect of building the pilot.
If you cannot state what an unacceptable output is, you are not ready to find out in production.
In regulated environments this is not a quality conversation. It is the delivery itself. The evaluation framework, the audit trail, the human oversight model: that is the product. The model is a component.
Compliance is not a phase
The standard pilot plan treats security, privacy, and regulatory review as a gate at the end. Build the thing, prove the value, then take it to compliance for approval.
This sequencing is how pilots die at month fourteen. The architecture that was fastest to demo is almost never the architecture that survives contact with data residency requirements, retention policies, access controls, and auditability. Retrofitting those constraints into a system that was built without them usually costs more than rebuilding, and by the time that becomes clear, the budget and the momentum are gone.
The teams that ship AI in regulated industries do the opposite. They treat the constraints as design inputs from day one. Where the data can live, what can leave the boundary, what has to be logged, where a human has to be in the loop. The pilot is built inside the constraints, which makes the pilot slightly slower and the production path radically shorter.
Asking permission at the end is not a compliance strategy. It is a bet that the reviewers will not look closely. In healthcare, in banking, in anything with a regulator, that bet loses.
What production-grade delivery looks like
The AI programs that reach production share five properties, and none of them is a better model.
An owner with authority. One person accountable for the path from pilot to production, with reach across the model, the data, the infrastructure, and the compliance boundary. Not a committee. Not a working group. A name.
Evaluation criteria written before the build. The definition of acceptable and unacceptable output exists before the first sprint, because it shapes the architecture, the oversight model, and the go-live decision. If the criteria are written after the system, they will be written to fit the system.
Compliance in the architecture. The regulatory constraints are design inputs, not a final gate. The system that reaches the review was built for the review.
An integration path into the workflow. Nobody uses a tool that lives outside their working environment, no matter how good the demo was. If the plan does not specify where the system sits inside the existing workflow, who acts on its output, and what they stop doing because of it, adoption is an assumption, not a plan.
A governance cadence that decides. Scale criteria and kill criteria, defined in advance, reviewed on a schedule. A pilot without a kill criterion never ends. It fades, and it takes budget and credibility with it as it goes.
The pilots that reach production are not the ones with the best models. They are the ones where someone owned the path.
The pilot was the easy part
A pilot answers the question of whether the technology can work. Production answers the question of whether the organization can run it. The second question is harder, it is organizational rather than technical, and it is the one that almost nobody staffs.
If your AI pilot is more than six months old and nobody can name the production owner, the evaluation criteria, and the compliance path, the pilot is not stuck. It is over. Nobody has said so yet.
The good news is that the diagnosis is also the treatment plan. Name the owner. Write the criteria. Put the constraints in the design. Define what scaling and stopping look like. None of this requires better technology. All of it requires the kind of delivery discipline that has always separated programs that finish from programs that fade.
The model was ready months ago. The question was never the model.