Case study · Custom AI Systems
Rebuilding a working AI demo into a system that survives real input
The demo worked. It answered cleanly, on tidy input, with a person watching and ready to retry. Then it had to run on real input, at real volume, with no one watching, and the gap showed. The prototype proved the idea. It was not yet a system.
- 01Validate input
- 02Guard and permit
- 03Run with fallbacks
- 04Observe and bound
- 05Deploy and reverse
- Focus
- Teams with a working AI prototype that has to run in production
- Problem
- A prototype that works in a demo on clean input breaks on real input, real volume, and no one watching.
- Approach
- Rebuild it as a layered system with a graded test set, live observability, security, failure handling, repeatable deployment, and bounded cost.
The problem
The prototype did its job. Someone built it to prove an idea, it answered well on the inputs it was shown, and the demo convinced the room. That is exactly what a prototype is for. The trouble starts when the same code is asked to run as a product: against input no one curated, at a volume no one sat and watched, with the retry button gone and a real user on the other end.
What worked in the demo was a single script holding the whole idea in one place: a prompt, a model call, and a happy path. That shape is fine for proving a concept. It has nowhere to put the parts a demo never needs, retries, timeouts, permission checks, a record of what happened, a way to roll back, and so it does none of them. The prototype is not a smaller version of the system. It is a different thing.
The reality behind the problem
Hardening a prototype is not a pass of cleanup on working code. The demo succeeded by avoiding every hard case: input arrived clean, the model answered on the first try, the one user was patient, and a human stood by to spot a wrong answer and run it again. Production removes all four of those supports at once. Input is malformed, adversarial, or empty. Calls time out and have to be retried without duplicating work. The user is gone, so a wrong answer ships unseen.
The honest read is that the prototype did the first 20 percent, the part that proves the idea is possible. The remaining 80 percent is the load-bearing work it skipped: validating untrusted input, retrying and degrading gracefully when a dependency fails, enforcing who is allowed to do what, measuring quality instead of eyeballing it, seeing inside the system when it runs, and being able to deploy and reverse a change safely. None of that was missing by mistake. A demo genuinely does not need it. A system cannot run without it.
What was assumed
The prototype works, so it is almost done.
What Stallwart asked
What happens to it on input no one curated, when the call fails, and when no one is watching?
The desired outcome
Before any code was reused, the engagement defined what the prototype had to become to run in production unattended. The criteria read as a production-readiness checklist.
- Run unattended on real, uncurated input without a person standing by to catch bad output.
- Treat every input as untrusted: validate it, and resist prompt injection, leaked secrets, and actions past a permission boundary.
- Fail safely when a model call times out or a dependency is down: retry, degrade, or fall back instead of breaking or hanging.
- Be measured against a graded test set, so a change to the prompt or model is checked before it ships, not discovered live.
- Be observable in production: every input, decision, and failure visible after the fact, not a black box.
- Deploy repeatably and roll back on demand, with token, latency, and spend held inside a known budget.
The system we designed
Stallwart rebuilds the prototype into a layered system while keeping the idea that made the demo work. The single script becomes a set of services with clear boundaries: an input layer that validates and sanitizes before anything reaches the model, the model interaction itself, an orchestration layer that handles retries, timeouts, and fallbacks, and a governance layer that enforces permissions and records every decision. The prompt that proved the concept is preserved. What changes is everything around it that a demo never needed.
The system is built to run without a person watching, which means it has to be honest about failure rather than optimistic about success. It validates what comes in, bounds what each request can cost, falls back to a safe response when a dependency is down, and writes a trail of what it did and why. That is the standard Stallwart holds on every build: a system that fails safely, reports what it is doing, and can be reversed when a change goes wrong.
How it works
- 01
Map what the demo skipped
The first pass is not rewriting, it is naming every assumption the prototype leaned on: clean input, a patient user, a dependency that never fails, a human ready to retry. Each assumption that production removes becomes a piece of work the system has to carry.
- 02
Split the script into layers
The single happy path is broken into services with boundaries: input validation, model interaction, orchestration with retries and fallbacks, and governance with permissions and logging. A guarantee made in one layer, that input is validated, that an action is permitted, then holds for everything downstream.
- 03
Grade, observe, and bound
A graded test set replaces eyeballing, so quality is measured before a change ships. Observability makes inputs, decisions, and failures visible in production. Token, latency, and spend limits are set per request, so cost cannot run away under real volume.
How it works
From a one-script demo to a layered system that runs unattended.
- 01
Validate input
Untrusted by default: checked and sanitized
- 02
Guard and permit
Injection resisted, permission boundary enforced
- 03
Run with fallbacks
Retries, timeouts, graceful degradation
- 04
Observe and bound
Decisions logged, token and spend capped
- 05
Deploy and reverse
Repeatable release, rollback on demand
Each stage is a step the prototype skipped because a demo did not need it, and a system cannot run without.
Architecture
What sits underneath.
The rebuilt prototype runs on the same four layers as every Stallwart build, so a guarantee made in one layer holds across the whole system rather than living in one unguarded script.
Intelligence
The model interaction the prototype got right is kept, with the prompt and model choice preserved, but it is now one bounded component rather than the whole program, called behind a validated input and a budget rather than directly from raw user text.
Orchestration
Retries with backoff, timeouts, graceful degradation, and fallbacks sit around the model call, so a slow or failed dependency produces a safe response instead of a hang or a crash, and work is not duplicated when a request is retried.
Governance
Input is treated as untrusted and validated before it reaches the model, prompt injection is resisted, secrets are kept out of prompts and logs, and permission checks decide what each request is allowed to do before it does it.
Production
Every input, decision, and failure is logged for observability, releases are repeatable and reversible through infrastructure as code, and token, latency, and spend are held inside a declared budget rather than left to float.
The hard parts
Where the engineering judgment was.
Keep the prototype's idea, replace its shape
The demo proved the approach works, and that is worth preserving. What fails in production is the single-script shape, not the idea. Rebuilding the structure while keeping the proven prompt and model is faster and safer than starting over or than patching the script.
The trade-off
The code that shipped the demo is mostly not the code that ships to production, which can feel like throwing away working software. It is not: the idea survives, the structure that could not scale does not.
Treat all input as untrusted from the first layer
A demo runs on input the builder chose. Production runs on input a stranger sends, which may be malformed, adversarial, or an attempt at prompt injection. Validating and sanitizing before the model ever sees the text is what stops a crafted input from becoming an unintended action.
The trade-off
Input validation and injection defenses add a layer the prototype did not have and reject some inputs a lenient demo would have accepted, a deliberate bias toward safety over permissiveness.
Measure quality with a graded set, not an eye
In a demo, a person reads the output and judges it. That does not scale and does not survive a prompt or model change. A graded test set of real cases turns quality into something checked automatically before a change ships, instead of a regression found by a user.
The trade-off
Building and maintaining the graded set is real work that produces no visible feature, spent on purpose because it is what makes every later change safe to make.
Reliability & guardrails
How it avoids the wrong call.
The guardrails exist so the system fails in a safe, visible way instead of a silent, confident one.
Input validation and injection defense
Every input is checked and sanitized before it reaches the model, and prompt injection is treated as an expected attack rather than an edge case, so a crafted input cannot steer the system into an action it should not take.
Retries, timeouts, and graceful degradation
Model and dependency calls run under timeouts and bounded retries, and when something stays down the system degrades to a safe, defined response instead of hanging or crashing, so one failed dependency does not take the whole request with it.
Secrets and permission boundaries
Secrets are kept out of prompts and logs, and permission checks run before an action, so the system cannot leak a credential into a response or act past the boundary of what the caller is allowed to do.
Bounded cost per request
Token count, latency, and spend are capped per request, so a pathological input or a retry storm cannot run the bill or the response time away under real volume.
Productionization
What turns it from a demo into software.
What separates a convincing demo from software a team can run unattended every day.
Evaluation harness
A graded set of real cases runs against the system, so a change to the prompt, model, or logic is scored against known-good output before it ships, and a regression is caught in review rather than in production.
Observability
Each request carries its input, the decisions it made, and any failure into the logs, so when an answer is wrong the team can see which input and which step produced it instead of guessing at a black box.
Repeatable, reversible deployment
Infrastructure as code makes releases repeatable and rollback a single deliberate action, so a bad change can be reversed in minutes rather than hot-fixed under pressure.
Cost and latency control
Token budgets, model choice, and latency limits are tuned and enforced, so per-request cost and response time stay inside a known envelope as volume grows rather than drifting upward unseen.
What changed
The operational change is that the system runs on its own. Instead of a demo that works when the input is clean and a person is watching, there is software that takes real, uncurated input, validates it, fails safely when a dependency is down, and records what it did. The prototype proved the idea could work once. The system does the same job reliably, unattended, at volume.
Because it is observable and reversible, it is also maintainable. When something goes wrong the team can see where, a change is scored before it ships, and a bad release can be rolled back. The result is not a flashier demo. It is the quiet shift from a thing that works in the room to a thing that keeps working out of it.
- The system runs on real input without a person standing by to catch bad output.
- A failed or slow dependency produces a safe, defined response instead of a crash or a hang.
- A prompt or model change is scored against a graded set before it ships, not after it breaks.
- When an answer is wrong the team can see which input and step caused it, and can roll the change back.
Ownership & handover
What you receive, and keep.
At handover, the hardened system is yours to run and extend, with nothing held back.
- Source code for the layered system: input validation, model interaction, orchestration, and governance.
- Infrastructure as code for repeatable deployment and one-step rollback.
- The graded evaluation set and harness, so you can keep scoring changes as the system grows.
- Observability dashboards and logs, with the trail of inputs, decisions, and failures already wired in.
- Runbooks for deploying, rolling back, tuning cost and latency, and extending the system safely.
What we learned
The lesson that generalizes: a prototype and a production system are different things, not different sizes of the same thing. A demo earns the first 20 percent, proving the idea is possible. The last 80 percent, validation, failure handling, security, evaluation, observability, and reversible deployment, is the load-bearing work a demo can skip and a system cannot. Rebuilding for it is not polishing the prototype. It is building the system the prototype only pointed at.
This work connects to
- Custom AI Systems
- Production AI systems
- AI engineering and hardening
- Evaluation and observability
- Reliability and failure handling
Frequently asked
Why does an AI prototype that works in a demo break in production?
A demo succeeds by avoiding every hard case: input is clean, the model answers on the first try, and a person is watching to catch a wrong answer and retry. Production removes all of those at once. Input is malformed or adversarial, calls time out, and no one is watching, so a wrong answer ships unseen. The prototype is not a smaller system, it is a different thing that skips the parts a demo does not need.
What is the 'last 80 percent' of an AI system?
The first 20 percent is proving the idea works, which is what a prototype does. The last 80 percent is the load-bearing work it skips: validating untrusted input, retrying and degrading gracefully when a dependency fails, enforcing permissions, measuring quality with a graded test set, making the system observable, and being able to deploy and roll back safely. A demo genuinely does not need any of it. A production system cannot run without it.
Can you reuse our existing prototype, or does it start over?
The idea and the proven parts are kept, the single-script shape is replaced. The prompt and model choice that made the demo work are preserved and become one bounded component. What gets built around them is everything a demo never needed: input validation, retries and fallbacks, permission checks, logging, evaluation, and reversible deployment. The approach survives, the structure that could not scale does not.
How do you handle untrusted input and prompt injection?
Input is treated as untrusted from the first layer. It is validated and sanitized before it reaches the model, prompt injection is handled as an expected attack rather than an edge case, secrets are kept out of prompts and logs, and permission checks run before any action. That way a crafted input cannot steer the system into leaking a credential or acting past what the caller is allowed to do.
How do you keep cost and latency under control in production?
Token count, latency, and spend are capped per request and enforced, not left to float. Model choice and retry behavior are tuned to hold response time and per-request cost inside a known budget, so a pathological input or a retry storm cannot run the bill or the latency away as real volume arrives. The limits are set deliberately and watched through the same observability as everything else.
Do you have an AI prototype that works in the demo but not in production?
Bring the prototype. We rebuild it into a system that survives real input, fails safely, and runs without a person watching.
Last updated: October 6, 2026