The AI pilot to production checklist (what the 80% skip)
Most AI pilots stall after the demo, and the reason is rarely the model. Here is the checklist that separates a demo from a system that actually runs.
14 min readBy Arun Saravanan
The checklist exists because the pilot proves only the easy 20 percent
A working pilot is not a signal that production will work. It is a signal that the interesting part of the problem is tractable, which was almost always the easy question. The actual production problem is the load-bearing 80 percent the pilot was allowed to skip: malformed inputs, retries, permissions, cost ceilings, rollback, audit trails, and the request that fits no category your scope anticipated. The checklist below exists so a working prototype gets judged on what still has to be built, not on how good the demo looked.
The failure rate is not a secret. The MIT NANDA 'State of AI in Business' report from August 2025, picked up by Fortune and the Yahoo Finance newsroom, found that roughly 95 percent of enterprise generative AI pilots are not producing measurable P&L impact. That number is not a model problem. It is a system-around-the-model problem, which is why swapping Claude for GPT or GPT for Gemini never fixes it. The pilot ends, nobody built the 80 percent, and the thing stalls at the handoff.
This post is the checklist your team should run against a working prototype before anyone writes a production budget. It is dual-use: engineers can treat the twelve items below as a build backlog, operators can treat them as a go or no-go gate. The structural reasons behind the gap are covered at length in [Why AI pilots fail to reach production](/blog/why-ai-pilots-dont-reach-production). This piece is the operational companion.
What the 95 percent number is actually counting
The MIT finding is sharper than the headline suggests. The report distinguishes two groups of failures: pilots that reached limited production and still did not change a P&L number, and pilots that never left the sandbox. Both are counted as failures, but they fail for different reasons. The first group shipped code that could not be trusted on the output side. The second group shipped code that could not be deployed on the operations side. A real production checklist has to cover both.
The gap between a working demo and a shippable system is roughly the gap between software that answers one question and software that answers every question the user will type in next. In the demo the engineer picks the input. In production the input is whatever a confused customer types into a text box at 2am, with a trailing attachment that no longer exists, in a language the pilot's eval set never saw. Treating the environment of production as part of the design, not a later concern, is the move that changes outcomes.
That is also why most 'MLOps for LLMs' checklists feel adjacent but miss. They borrow their vocabulary from classical ML, where retraining and feature drift were the main worries. In an LLM system the model is a rented commodity you do not retrain. The real failure modes are prompt injection, retrieval drift, context window overrun, cost spikes from a retry loop, and silent output degradation nobody notices because there is no eval running. The checklist has to be written for that reality.
In plain terms for the non-engineer reading along: the headline number does not mean the models are bad. It means the engineering work around the models is almost never scoped into the pilot budget, so the pilot ends with the easy part done and the hard part untouched. The checklist below is what the hard part looks like, item by item.
The twelve items that separate a demo from a system
Each item below is a yes/no gate. If the answer is no, it is a build item, not a footnote. Any pilot being promoted to production should pass all twelve. The order matters: the first six are what keeps the system from doing damage on a bad day. The last six are what keeps it operable on an average one.
- Input validation and schema guardrails. Every incoming input is schema-checked, length-bounded, and sanity-screened before it reaches a model. Prompt injection attempts, oversized payloads, and malformed JSON fail fast with a logged reason, not a surprising model response. The failure mode this prevents: an attacker pastes a crafted string into a free-text field and the model reveals a system prompt, a tool definition, or another user's data.
- Grounded output constraints. Where the output must match a shape (JSON, a document type, a yes/no decision), the model is constrained to produce that shape and a parser verifies it before the output is used. Free-text responses pass through a classifier that flags the categories the business decided were unsafe to ship. The failure mode this prevents: a confidently wrong output reaches a downstream system that trusts it and acts on it.
- Timeouts, bounded retries, and defined fallbacks. Every model and third-party call has a defined timeout, a capped retry policy (not infinite), and a defined fallback: a cheaper model, a cached response, or a graceful decline. A dependency that stops responding does not take the whole workflow with it. The failure mode this prevents: a model API has a bad hour and every request in your queue retries forever, burning cost and jamming the pipeline.
- Permissions and data boundaries. The system can only read and act within explicit boundaries. A customer in tenant A cannot see data from tenant B through a prompt. Service accounts hold the minimum role required to do the job. Secrets live in a vault, not in prompts. The failure mode this prevents: cross-tenant data leakage, which is the single most common AI incident in multi-tenant SaaS and the one that ends enterprise contracts.
- Human escalation paths. Classes of request that are high-stakes, out-of-category, or low-confidence route to a person by design. The handoff carries enough context that the human can act in under a minute, not reconstruct the case from scratch. The failure mode this prevents: the system quietly handles the request it should have escalated and nobody notices until the customer calls.
- Rate limits and cost ceilings. Per-user, per-tenant, and global rate limits exist and are enforced. A circuit breaker caps spend per hour. A retry loop or a traffic spike cannot produce a surprise invoice. The failure mode this prevents: a bug or a bot produces a five-figure overnight bill the finance team learns about at month-end.
- Request-level logging and tracing. Every request produces a structured log: input, model, prompt version, output, tokens in and out, cost, latency, and which guardrails fired. Traces thread through retrieval, model calls, and post-processing. You can answer 'what did the system do for this specific user at this specific time' in under 60 seconds. The failure mode this prevents: a customer complaint you cannot investigate because nobody logged enough to reconstruct what happened.
- Continuous evaluation. A golden set of input/output pairs, grown from real production traffic, runs on every change. Drift is measured, not assumed to be absent. Hallucination, task accuracy, and format compliance each have thresholds, and a build that drops below any of them does not ship. The failure mode this prevents: quality rots silently over months as prompts, retrieval indexes, and model vendors shift under the system.
- Rollback and versioning. Model, prompt, retrieval index, and application code are each versioned independently. Any one can be reverted in minutes without reverting the others. 'Rollback' is a documented procedure a person on-call can run at 3am, not a Slack thread about who deployed what. The failure mode this prevents: a bad change means a day-long incident instead of a two-minute revert.
- Observability dashboards and alerting. Latency, error rate, cost per request, and eval scores sit on a dashboard. On-call gets paged when any goes outside its agreed band. The dashboard answers 'is it working right now' without a person logging in to find out. The failure mode this prevents: the system is broken for a quiet eight hours and the team learns about it from a customer.
- Named ownership. One person is paged on failure and empowered to act. The eval set, the prompt library, and the model config each have a named owner who approves changes. 'Shared ownership' means no owner, and no-owner is where production AI systems slowly rot. The failure mode this prevents: a system that worked at launch and nobody has looked at in six months.
- A written runbook. A new engineer takes the page at 3am, opens one document, and acts. The runbook lists the top five known failure modes, how to roll back each component, how to escalate to the vendor, and how to communicate downtime to users. The failure mode this prevents: a system that only one person knows how to recover, and that person is on vacation.
The evaluation and observability layer most pilots skip
Items seven through ten are where most pilots fail silently, which is the worst kind of failure. A system without continuous evaluation drifts. A prompt change that improves the demo subtly worsens an edge case nobody re-tested. A retrieval index that was accurate at launch degrades as the underlying content changes. A model vendor updates the model behind an API and the behavior shifts. Nothing visibly breaks. The numbers move the wrong way, slowly, and the team hears about it from a customer six weeks later.
The fix is to make evaluation a gating step, not a dashboard. Build the golden set from real production traffic, not synthetic examples. Grow it every time a new failure mode is caught, because a failure that was never written into a test will happen again. Run the eval on every change to prompt, model, or retrieval index, and gate deploys on the result the same way you gate on unit tests. A deploy that drops hallucination accuracy below the agreed threshold does not merge, regardless of how green the rest of CI looks. For the longer version of this argument, see [How to evaluate AI systems](/blog/how-to-evaluate-ai-systems).
Observability sits one layer up. The question observability answers is not 'is the model correct' (that is evaluation) but 'is the system healthy and affordable right now'. Latency, error rate, token spend, which prompt version is live, which model is live, how many requests hit the fallback path. These are production SRE metrics applied to an LLM-backed workflow. If your pilot has none of this, you do not have a system to promote; you have a notebook that happens to run.
In plain terms: evaluation measures whether the AI is giving the right answers. Observability measures whether it is still running and how much it is costing. You need both. Teams that skip evaluation discover quality problems from angry customers. Teams that skip observability discover cost problems from surprise invoices. Teams that skip both end up in the 95 percent.
The five go or no-go questions before you fund production
The twelve-item checklist is for the build team. The five questions below are for the person signing the budget. Any pilot that cannot answer all five with specifics is being promoted on hope, not evidence. These should take ten minutes in a room, not a slide deck.
- What does it do when the input is wrong? Not 'should' do. What does the current code do, right now, on a malformed, adversarial, or out-of-category input. If the answer is 'we have not tested that', the pilot is not ready for the production budget conversation.
- Who is paged when it fails, and what can they do at 2am? If the answer is a shared mailbox, nobody is paged. If the answer is 'we would roll back', ask who has done that end-to-end before, and how long it took them.
- How do you turn it off in isolation, without turning off everything around it? A system that can only be reverted by reverting the whole product is a system that will not be reverted when the time comes, because the cost of the revert is too high.
- How do you know it is still correct next month, not just correct in the demo? The answer is an eval set and a cadence. 'We will keep an eye on it' is not an answer.
- What does one unit of work cost, and what stops that cost from running away? Both numbers. If the team cannot give you a cost per request within a factor of two, the business case is still a sketch.
Where this checklist still fails you
The checklist cannot tell you whether the use case was the right use case. A system that passes every item and automates something that did not need automating is still a bad investment. The decision on the use case happens before the checklist, when a technical lead and an operations lead decide what specifically changes for a named human when the system works. For how to scope a pilot so the production question is answerable at the end of it, see [What 'pilot with guardrails' actually means](/blog/pilot-with-guardrails-explained).
The checklist is also a floor, not a ceiling. A high-stakes application (anything touching money, health, legal, or safety) adds items around auditability, bias testing, model documentation, data lineage, and regulatory evidence. The twelve items here are the minimum any production AI system should pass. If you are building in a regulated vertical, treat this as the first of two checklists you need to pass.
And the checklist does not substitute for ownership. A system can pass every gate on launch day and rot six months later because no engineering team was handed it. Production AI is operated, not shipped. Who owns the eval set a year from now is a harder question than any item above, and it is also the one that most reliably decides whether the system survives its first on-call rotation.
What to do before the next steering meeting
Run the twelve items against your current pilot before the next steering meeting. Score each as yes, no, or partial. Any no is a work item. Any partial is a work item that has been under-estimated. The output is a specific list of what has to be built, costed, and sequenced before go-live, which is a much more useful conversation to walk into than 'is the pilot ready'.
If you want a second set of eyes on a specific build, Stallwart does this work: production AI engineering for US B2B teams moving from a working demo to a system that stays up. We plug into an existing pilot, run the gate, and come back with the list, the sequencing, and the build plan. The deliverable is a system that passes the checklist on launch day and still passes it a year later, with the owners named.
The short version
- A pilot proves the interesting 20 percent of a problem. Production is the load-bearing 80 percent, and most pilots are not scoped to build it.
- The AI production checklist has twelve items in two blocks: six that keep the system from doing damage on a bad day, six that keep it operable on an average one.
- Evaluation and observability are different problems. Evaluation measures whether answers are correct; observability measures whether the system is healthy and affordable.
- The five go or no-go gate questions a budget owner should answer before funding production cover input failure handling, on-call, isolated rollback, continuous correctness, and cost ceiling.
- A model swap almost never fixes a production AI failure. The failure is in the system around the model, not the model.
- Named ownership is the item most reliably correlated with a production AI system surviving its first year.
Questions this raises
- What is an AI production readiness checklist?
- It is a specific list of system-level items a working AI prototype has to pass before deployment: input validation, grounded output constraints, timeouts and fallbacks, permissions, human escalation paths, rate and cost limits, request-level logging, continuous evaluation, rollback and versioning, observability, named ownership, and a written runbook. Each item has a yes/no answer and a documented failure mode. If any item is no, it is a build item, not a footnote.
- Why do most AI pilots fail to reach production?
- Because the pilot was scoped to prove the model works on clean input, and production is a system-around-the-model problem the pilot was never scoped to build. The MIT NANDA 'State of AI in Business' report from August 2025 found roughly 95 percent of enterprise generative AI pilots are not producing measurable P&L impact. The common cause is not the model; it is missing evaluation, observability, rollback, cost control, and named ownership.
- What percentage of enterprise AI projects fail to reach production?
- The widely cited figure is around 95 percent, from the MIT NANDA 'State of AI in Business' report published in August 2025 and covered by Fortune and other outlets. Other industry surveys put the figure in the 80 to 90 percent range depending on how 'production' is defined. The common factor across the surveys is that the gap is not model quality; it is the engineering work around the model that pilots do not include.
- What is the hardest part of putting an LLM in production?
- It is almost never the model. The hardest parts are continuous evaluation that catches silent output drift, cross-tenant data isolation, cost control under retry loops, and isolated rollback. These problems do not exist in the demo because the demo runs on clean input chosen by the engineer. They surface the first week after launch, when real users send input the pilot never saw.
- How do you evaluate an AI system in production?
- Build a golden set of input/output pairs from real production traffic, grow it every time a new failure mode is caught, and run it on every change to prompt, model, or retrieval index. Treat the eval score as a merge gate: a change that drops hallucination or task accuracy below the agreed threshold does not ship. Dashboards are useful for triage, but the gate on deploys is what prevents quality from rotting silently.
- How do you put a cost ceiling on a production AI system?
- Enforce three layers: per-user and per-tenant rate limits in the application, bounded retries with capped backoff on every model call, and a global circuit breaker that halts the system when hourly spend exceeds a threshold. Log cost per request and alert on cost per request moving outside its band, which catches silent model switches or retrieval regressions before they show up in the monthly invoice.
- Who should own a production AI system after it ships?
- A named person who is paged on failure and empowered to act, with the eval set, the prompt library, and the model config each having a named owner who approves changes. Shared ownership is where production AI quietly rots: nobody updates the eval set, nobody re-tests after a vendor model update, nobody notices the cost drift. Ownership is not a role on an org chart; it is a page in a runbook with a name on it.
Recognize this in your own operation?
Bring us the version of it happening in your business and we will tell you which part a system can take over.