Case study · Custom AI Systems
An AI agent that knows when not to act
Giving an agent the ability to act is the easy half. The hard half is restraint: deciding which actions it must never take alone, calibrating when it is unsure, and routing those moments to a person. The interesting question was never 'can it act.' It was 'when should it stop, ask, or escalate.'
- 01Request and context
- 02Plan and classify
- 03Confidence gate
- 04Approval on irreversible
- 05Act and log
- Focus
- Operations and back-office teams running agentic workflows in B2B software and services
- Problem
- An agent that can act will also act when it should not, confidently and irreversibly, unless restraint is built in.
- Approach
- Guardrails as engineering: confidence thresholds, hard tool limits, approval gates on irreversible actions, and a full audit trail.
The problem
An agent that can call tools can also call the wrong one. It can send the message, close the ticket, issue the refund, or change the record, and it will do so with the same fluent confidence whether it is right or wrong. The capability to act arrives first. The judgment about when not to act does not come with it.
A demo shows the agent completing a task end to end. Production shows the other cases: the ambiguous request, the thin context, the irreversible step taken on a guess. The hard part is not teaching the agent to do more. It is engineering the agent to stop, ask, or escalate at the exact points where acting alone would be a mistake.
The reality behind the problem
The obvious fix, tell the agent to be careful in its instructions, does not hold under load. A prompt that asks for caution is a suggestion, not a control. The agent still decides, in the moment, with whatever context it has, and nothing structural stops it from taking an irreversible action on a weak inference.
The failure mode is quiet and expensive. An agent rarely stalls visibly. It proceeds, picks a plausible action, and executes it. Without a measured sense of its own confidence and hard limits on what it may do unattended, the agent is most dangerous precisely when it is least sure, because its output reads the same either way.
What was assumed
We need an agent that can do everything a person can do.
What Stallwart asked
Which actions must never be taken on a guess, and how does the agent know it is on one?
The desired outcome
Before a model or a tool set was chosen, the engagement defined what the agent had to do to be trusted to run with real authority.
- Classify actions by reversibility, and gate the irreversible ones behind an approval a person gives.
- Carry a calibrated confidence signal, so 'unsure' is detected and acted on rather than hidden.
- Escalate to a person in a form they can act on, with the context and the proposed action attached.
- Hold hard limits on tool use that no instruction in the moment can widen.
- Log every decision, the inputs behind it, and whether it was taken or escalated, so the trail is queryable after the fact.
- Stay reversible where the engagement requires it, so a taken action can be undone or rolled back.
The system we designed
Stallwart builds an agent whose restraint is part of the system, not part of the prompt. Every tool the agent can call is classified by what it changes and whether that change can be undone. Reversible, low-stakes actions run on the agent's own judgment above a confidence threshold. Irreversible or high-stakes actions sit behind an approval gate: the agent proposes, a person confirms, and only then does the action execute. The thresholds and the gates are runtime controls, enforced in code, not guidance the model can talk itself past.
The agent is built to report when it is unsure and to escalate rather than guess. When its confidence falls below the line, or a request touches an action it is not cleared to take alone, it stops and hands the decision to a person with the full context attached. That is the standard Stallwart holds across every build: a system that knows what it is running, explains itself, escalates what it should not decide, and is reversible.
How it works
- 01
Map the actions before the agent runs
Every tool the agent can reach is classified by what it changes and whether it can be undone. That classification, not a line in the prompt, decides which actions run on the agent's judgment and which wait for a person. The governance is set here, before a single request is handled.
- 02
Score confidence, then route
Each proposed action carries a confidence signal built from the quality of its context and its inputs. Above the threshold on a reversible action, the agent proceeds. Below it, or on an irreversible action, the agent stops and escalates rather than acting on a weak inference.
- 03
Gate, act, and record
Irreversible and high-stakes actions pass through an approval gate a person controls. When an action is taken, it is taken reversibly where the engagement requires it, and every decision, its inputs, and its outcome are written to an audit trail that stays queryable.
How it works
From a request to an action taken, approved, or escalated.
- 01
Request and context
What is being asked, with what the agent can verify
- 02
Plan and classify
Proposed action scored for stakes and reversibility
- 03
Confidence gate
Below the line, the agent escalates instead of acting
- 04
Approval on irreversible
High-stakes steps wait for a person to confirm
- 05
Act and log
Action taken, reversibly where required, and recorded
Each stage either clears the agent to act, holds it for approval, or routes the decision to a person, with the whole path logged.
Architecture
What sits underneath.
Underneath the workflow, the agent runs on the same four layers as every Stallwart build, so a guarantee made in one layer holds across the whole system.
Intelligence
The model plans a response to the request and proposes an action, but it also produces a confidence signal drawn from the strength of its context and inputs. That signal is a first-class output, carried forward so the system can act on how sure the agent is, not just on what it decided.
Orchestration
Tools are registered with a stakes and reversibility class, and the orchestrator routes each proposed action by it: run now, hold for approval, or escalate. Hard limits on which tools the agent may call, and how far, are enforced here, where no instruction in the moment can widen them.
Governance
Policy lives as runtime controls, not prose. Confidence thresholds, approval gates, and tool limits are configuration the system enforces on every decision, and every request, the action proposed, the confidence, and whether it was taken or escalated are written to an audit trail that stays queryable after the fact.
Production
Actions are designed to be reversible where the engagement requires it, with the path to undo or roll back one built in rather than bolted on. Observability on the live agent surfaces what it is doing, what it is escalating, and where it is stopping, so the system's behavior is visible in operation.
The hard parts
Where the engineering judgment was.
Reversibility decides the gate, not importance
The cleanest line for what the agent may do alone is whether the action can be undone. A reversible action taken on a slightly wrong inference is a correction. An irreversible one is a loss. Classifying by reversibility, rather than by a vague sense of how important a task feels, gives the gate a rule that holds under real traffic.
The trade-off
Some reversible actions carry real weight, and some irreversible ones are trivial. The classification is reviewed per tool against real cases rather than assumed, which takes judgment up front instead of a default.
Calibrate confidence so 'unsure' actually fires
A confidence score is only useful if the threshold is set where the agent's uncertainty maps to real risk. The valuable behavior is the agent escalating when it is on a guess, so the signal has to be calibrated against cases where it was wrong, not left at a number that looks reasonable.
The trade-off
A threshold tuned toward caution means the agent escalates some questions a person would have let it answer. That bias toward asking over acting is deliberate, and it is the price of trusting the agent with authority at all.
Design escalation a person can actually use
An escalation that arrives as a vague alert is ignored, and an agent whose escalations are ignored is back to acting unsupervised. The handoff has to carry the request, the proposed action, and the context behind it, in a form a person can approve or redirect in seconds.
The trade-off
Building the escalation path well is more work than firing a notification, and it couples the agent to a human review surface. That coupling is the point: the agent is only as safe as the moment a person can step in.
Reliability & guardrails
How it avoids the wrong call.
The guardrails exist to make the restrained failure mode the default one.
Approval gates on irreversible actions
Actions that cannot be undone wait for a person to confirm before they execute. The agent proposes and holds; it does not take an irreversible step on its own judgment, however confident it reports being.
Calibrated confidence thresholds
A confidence signal gates each action. Below the line, the agent escalates instead of acting, so a weak inference never turns into a taken action that reads as a sure one.
Hard tool limits
The set of tools the agent can call, and how far it can go with each, is bounded in code. No instruction supplied in the moment can widen that boundary, so the agent's reach is fixed rather than negotiable.
Audit trail behind every decision
Every request, the action proposed, the confidence behind it, and whether it was taken or escalated are logged. When the agent does something surprising, the trail shows why it decided what it did.
Productionization
What turns it from a demo into software.
What separates a convincing demo from an agent a team lets run against real systems.
Evaluation harness
A graded set of real scenarios runs continuously, including the cases where the right move is to stop or escalate, so a change to the prompt, model, or policy is measured against known-good behavior before it ships rather than discovered in production.
Reversibility and rollback
Actions are designed to be undone where the engagement requires it, and the path to roll back a taken action is built and tested, so a mistake is recoverable instead of permanent.
Observability on live behavior
The running agent exposes what it is doing, what it is holding for approval, and where it is escalating, so a shift in its behavior is visible in operation rather than inferred after something goes wrong.
Autonomy widened behind evidence
The agent starts with a narrow set of actions it can take alone. The boundary is widened only where the evaluation and the audit trail show it is safe to, so autonomy grows on evidence rather than on optimism.
What changed
The operational change is in what the agent does when it is unsure. Instead of proceeding on a guess, it stops: irreversible actions sit behind an approval a person gives, and low-confidence decisions route to a person with the context attached. The agent still does the routine work on its own. It just stops being the one to decide the moments it should not own.
Because the agent escalates, logs, and can be reversed, it earns the kind of trust an unsupervised one never does. A person can see what it did and why, step in before an irreversible action, and undo one that was taken. The most valuable behavior the system has is a well-timed stop.
- Irreversible actions sit behind an approval a person gives, instead of running on the agent's own judgment.
- The agent escalates when it is unsure, rather than taking a plausible action on a weak inference.
- Every decision is logged with its inputs, so the agent's behavior can be audited rather than guessed at.
- Taken actions are reversible where the engagement requires it, so a mistake is recoverable instead of permanent.
Ownership & handover
What you receive, and keep.
At handover, the agent is yours to run and extend, with nothing held back.
- Source code for the agent, its tool integrations, and the orchestration that routes each action.
- The policy configuration: the confidence thresholds, approval gates, and tool limits, as controls you can adjust.
- Infrastructure as code for the services, the queues, and the audit store behind the agent.
- The evaluation set and harness, including the scenarios where the correct behavior is to stop or escalate.
- Documentation of the action classification, the escalation path, the audit trail, and the reversibility model.
What we learned
The lesson that generalizes: an autonomous agent is a governance problem long before it is a capability problem. Making an agent act is the undifferentiated part. The value is in how actions are classified by reversibility, how confidence is calibrated so uncertainty fires, and how cleanly the agent stops and hands a decision to a person, which is exactly the work a demo skips and production demands.
This work connects to
- Custom AI Systems
- Autonomous agents
- Agent governance and guardrails
- Human-in-the-loop systems
- Production AI systems
Frequently asked
What makes an AI agent safe to run autonomously?
Restraint that is built into the system rather than requested in the prompt. A safe agent classifies its actions by whether they can be undone, gates the irreversible ones behind a person's approval, carries a calibrated confidence signal so it can tell when it is unsure, and logs every decision. The capability to act is the easy part; the engineering is in when the agent should stop, ask, or escalate.
Why gate an agent on reversibility instead of on how important an action is?
Because reversibility is a cleaner rule that holds under real traffic. A reversible action taken on a slightly wrong inference is a correction a person can fix; an irreversible one is a permanent loss. Classifying each tool by whether its effect can be undone gives the approval gate a concrete line, rather than relying on a vague sense of how important a task feels in the moment.
How does the agent know when it is unsure?
It carries a confidence signal built from the strength of its context and inputs, and that signal is calibrated against cases where the agent was actually wrong, not left at a number that merely looks reasonable. When confidence falls below the threshold, the agent escalates to a person instead of acting. Tuning the threshold toward caution means the agent sometimes asks about a question it could have answered, which is the deliberate price of trusting it with authority.
What is human-in-the-loop escalation, and how is it designed to be usable?
It is the path the agent takes when it should not decide alone: it stops and hands the decision to a person. To be usable, the handoff carries the original request, the action the agent proposes, and the context behind it, so a person can approve or redirect it in seconds. An escalation that arrives as a vague alert gets ignored, and an agent whose escalations are ignored is back to running unsupervised.
Can actions the agent takes be undone?
Systems are designed so actions are logged and reversible where the engagement requires it. Irreversible steps sit behind an approval gate before they execute, and for actions that can be undone, the path to roll one back is built and tested rather than assumed. Combined with an audit trail behind every decision, that means a mistake is recoverable and traceable instead of permanent and silent.
Do you need an agent that acts, or one that knows when not to?
Bring the workflow. We build an agent whose restraint is engineered: it acts where it is safe, stops where it is not, and keeps a record you can check.
Last updated: October 6, 2026