Case study · AI Agents & Automation
Running a multi-step back-office process as one system
The process already worked. It just ran as a chain of repetitive human steps spread across spreadsheets, email, dashboards, and a few internal APIs, and each step waited on a person to move it along. Nothing owned the whole run, nothing held its state, and the work stalled whenever someone was busy. The problem was never any single step. It was the handoffs between them.
- 01Intake
- 02Orchestrated steps
- 03Retry and idempotency
- 04Approval gate
- 05Completion and audit
- Focus
- Operations teams running multi-step internal processes across tools
- Problem
- A process ran as separate manual steps across tools, with no single owner, no shared record, and a stall whenever one person was unavailable.
- Approach
- One orchestrated system that runs the steps end to end, holds state across them, retries transient failures, and routes genuinely ambiguous decisions to a person for approval before continuing.
The problem
The process did its job, but it did it as a relay. One person pulled a record from a spreadsheet, checked it against a dashboard, sent an email, waited for a reply, updated a second sheet, then called an internal API by hand. Each step was simple on its own. The cost sat in the gaps between them, where a run waited for a person to notice it was their turn.
Because no single thing owned the process, its state lived in people's heads and in scattered tabs. There was no shared record of where a given run was, what had already happened to it, or what came next. When someone was out, their step simply paused, and the only way to know the status of anything was to ask the person holding it.
The reality behind the problem
The obvious fix, a script that fires the steps one after another, breaks the first time reality intervenes. A step calls an API that times out, an email never gets a reply, a record is malformed, and a naive script either halts in the middle or, worse, reruns and fires the same effect twice. Half-finished runs are the normal case in a real back office, not the exception.
The genuinely hard part is not automating any one step. It is holding the state of a run that spans hours or days, resuming it exactly where it left off after a failure, making sure an action that already happened is never repeated, and knowing which decisions a machine may take on its own versus which ones a person has to approve. That is orchestration, and it is where a demo quietly skips the work that production demands.
What was assumed
We need to automate each of these steps.
What Stallwart asked
What owns the run, holds its state, and decides what happens when a step fails halfway through?
The desired outcome
Before any step was automated, the engagement defined what the system had to guarantee to be trusted with a live process.
- Run the steps end to end without waiting on a person to move each one along.
- Hold the full state of every run, so its status is a fact the system knows, not something someone remembers.
- Retry transient failures on their own, and resume a run from the exact step it stopped at.
- Apply each external effect once, so a retry or a restart never sends the same email or posts the same record twice.
- Route genuinely ambiguous decisions to a person for approval, and continue only once they decide.
- Keep a single audit trail of every run, every step, and every approval, queryable after the fact.
The system we designed
Stallwart builds one orchestrated system that owns the process from intake to completion. A run becomes a durable object with explicit state: the system knows which step it is on, what each step produced, and what remains. It drives the sequence itself, calling the internal APIs, reading the sources, and advancing the record, so the work no longer depends on a person noticing it is their turn.
The system is built around failure rather than against it. Transient errors are retried with backoff, and because every external effect is made idempotent, a retry or a restart never double-fires. Where a decision is clear, the system takes it. Where judgment is required, it pauses at an approval gate, presents the case to a person, and resumes once they decide. Every run, step, retry, and approval lands in a single audit trail, so the scattered manual handoffs are replaced by one record of what happened and why.
How it works
- 01
Model the run, not the steps
The process is captured as an explicit state machine: the states a run can be in, the transitions between them, and the data each step reads and writes. The run becomes a durable object the system owns, so its status is always known rather than inferred from who last touched it.
- 02
Drive the sequence, survive the failures
The system executes each step, calling the internal APIs and advancing the record. Transient failures are retried with backoff, every external effect is guarded so it fires once, and a run that stops mid-sequence resumes from the exact step it reached rather than starting over.
- 03
Gate the judgment calls
A confidence signal separates the clear cases from the ambiguous ones. Clear cases proceed on their own. Ambiguous cases pause at an approval gate, where a person sees the full context and decides, and the run continues from there with that decision recorded.
How it works
From a relay of manual handoffs to one owned, resumable run.
- 01
Intake
A run starts as one durable record
- 02
Orchestrated steps
The system drives the sequence and holds state
- 03
Retry and idempotency
Transient failures retried, each effect applied once
- 04
Approval gate
Ambiguous decisions routed to a person
- 05
Completion and audit
Run closes, full trail retained
Each run is a durable object the system advances on its own, pausing only where a person's judgment is required.
Architecture
What sits underneath.
Underneath the workflow, the system runs on the same four layers as every Stallwart build, so a guarantee made in one layer holds across the whole process.
Intelligence
Classification and decisioning for each step: which path a run should take, and how confident that judgment is. The confidence signal is what decides whether a step proceeds on its own or stops at an approval gate for a person.
Orchestration
The durable state machine that owns each run. It sequences the steps, persists state across hours or days, retries transient failures with backoff, enforces idempotency so every external effect fires once, and resumes a failed run from the step it stopped at.
Governance
Approval gates, permissions, and the single audit trail. Every run, step, retry, and human decision is recorded with who approved what and when, so the scattered manual handoffs become one queryable record of the process.
Production
Observability across every run, rollback for a release that misbehaves, and an evaluation harness that checks changes to the flow before they ship. A run in flight is inspectable, and a bad change is reversible without losing in-progress work.
The hard parts
Where the engineering judgment was.
Make every external effect idempotent
In a process that retries and resumes, the same step will run more than once. Without idempotency, a retry sends a second email or posts a duplicate record, which is often worse than the original failure. Each effect is keyed so that repeating it is a safe no-op.
The trade-off
It means designing an idempotency key and a dedup check for every action that touches the outside world, which is more engineering per step than a script that assumes each one runs exactly once.
Persist state, do not hold it in memory
A back-office run can span hours or days and must survive a restart, a deploy, or a crash. State kept only in a running process disappears the moment it stops. Writing each transition to durable storage is what lets a run resume from exactly where it was.
The trade-off
Durable state adds storage and a write on every transition, chosen on purpose because a run that cannot be resumed is a run that silently loses work.
Place approval gates where judgment lives, not everywhere
Routing every step to a person recreates the original bottleneck. Routing none of them automates mistakes at speed. The gates sit precisely at the decisions where the context is genuinely ambiguous, so people spend their attention only where it changes the outcome.
The trade-off
Deciding where a gate belongs takes real work with the operators, and the boundary is tuned against real runs rather than set once from a diagram.
Reliability & guardrails
How it avoids the wrong call.
The guardrails exist so that a partial failure degrades into a safe, resumable state rather than a corrupted one.
Exactly-once effects
Every action that touches an external system is guarded by an idempotency key, so a retry or a restart never fires the same effect twice. Repeating a completed step is a safe no-op, not a duplicate.
Resumable runs
State is persisted at each transition, so a run that stops partway through, from a timeout, a crash, or a deploy, resumes from the exact step it reached instead of restarting or being abandoned.
Confidence-gated approval
A confidence signal gates each judgment call. Below the threshold, the run pauses and routes to a person rather than proceeding, so an ambiguous case never resolves itself with a guess.
Bounded retries with escalation
Transient failures are retried with backoff up to a limit. A step that keeps failing stops retrying and is surfaced for a person to look at, so a persistent fault is raised rather than hidden behind endless retries.
Productionization
What turns it from a demo into software.
What separates a convincing demo from a system a team runs its real process on every day.
Observability per run
Every run carries its current state, its step history, its retries, and its approvals into the logs, so when something goes wrong the team can see which run, which step, and which decision produced it.
Evaluation harness
A set of recorded runs and edge cases is replayed against any change to the flow, so a new path or a tuned gate is checked against known behavior before it ships rather than discovered on a live process.
Rollback and safe deploys
A release that misbehaves can be rolled back without losing runs already in flight, because state lives in durable storage rather than in the running version of the code.
Single audit trail
Every run, step, retry, and human approval is written to one trail, so the record of the process is a query rather than a reconstruction from email threads and spreadsheet history.
What changed
The operational change is in what owns the process. Instead of a run living in someone's tabs and advancing only when that person is free, the system owns it: the run proceeds unattended, holds its own state, and pauses only at the approval gates where a person's judgment is actually required. The process no longer stalls because one person is busy.
Because every run and every decision lands in one audit trail, the status of the work becomes a fact the system can answer rather than a question someone has to chase. The scattered handoffs across spreadsheets, email, and dashboards are replaced by a single record of what ran, what each step did, and who approved what.
- The process runs unattended except at approval points, instead of waiting on a person to move each step.
- A run's status is state the system holds, not something tracked across tabs and remembered.
- A step that fails is retried and resumed rather than quietly stalling or silently firing twice.
- The scattered manual handoffs become one audit trail of every run, step, and approval.
Ownership & handover
What you receive, and keep.
At handover, the system is yours to run and extend, with nothing held back.
- Source code for the orchestration engine, the step integrations, and the approval gates.
- Infrastructure as code for the state store, the services, and the deployment.
- The evaluation harness and recorded runs, so you can test changes to the flow as it grows.
- Runbooks for adding a step, adjusting a gate, resuming a stuck run, and rolling back a release.
- Documentation of the state machine, the idempotency model, the gate placement, and the audit trail.
What we learned
The lesson that generalizes: automating a back-office process is an orchestration problem long before it is an automation problem. Firing each step is the easy part. The value is in what owns the run, how its state survives a failure, how an effect is made to happen exactly once, and where a person's judgment is required. That is the work a demo skips and a live process demands, and it is what turns a chain of manual handoffs into one system a team can rely on.
This work connects to
- AI Agents & Automation
- Workflow orchestration
- Human-in-the-loop approval
- Back-office automation
- Production AI systems
Frequently asked
What is workflow orchestration for a back-office process?
Workflow orchestration means one system owns a multi-step process from start to finish instead of people moving it along by hand. It holds the state of each run, drives the steps in order, calls the internal tools and APIs, retries failures, and pauses for a person only where a decision needs human judgment. It replaces a relay of manual handoffs with a single owned, resumable run.
How is this different from a script that runs the steps in sequence?
A plain script assumes every step succeeds and runs once. A real process fails halfway, waits on a reply, or restarts mid-run. Orchestration adds durable state so a run can be resumed from where it stopped, idempotency so a retry never fires the same effect twice, bounded retries with escalation, and approval gates for ambiguous decisions. The difference is how it behaves when a step fails, which is the normal case, not the exception.
Where do people still make decisions in the process?
At approval gates placed precisely where judgment is required. A confidence signal separates the clear cases, which the system handles on its own, from the ambiguous ones, which pause and route to a person with the full context. The person decides, the decision is recorded, and the run continues. People spend their attention only on the decisions that change the outcome.
How does the system avoid firing the same action twice?
Every action that touches an external system is guarded by an idempotency key, so repeating a step that already completed is a safe no-op rather than a duplicate. Combined with durable state that records which steps are done, this means a retry after a timeout, or a restart after a deploy, never sends a second email or posts a duplicate record.
What happens to a run if a step fails or the system restarts?
State is persisted at every transition, so a run that stops partway through resumes from the exact step it reached instead of restarting or being lost. Transient failures are retried with backoff up to a limit, and a step that keeps failing stops retrying and is surfaced for a person rather than hidden. A release can be rolled back without losing runs already in flight.
Does your process stall every time it reaches a busy person's inbox?
Bring the steps and the tools they touch. We build one system that owns the run, holds its state, and asks a person only where judgment is required.
Last updated: October 6, 2026