Skip to content

Case study · AI Infrastructure & RAG

From company data to a business action, behind a human gate

Answering a question is the easy half. The work that moves the business is the step after: updating the record, routing the decision, flagging the exception, triggering the next task. That step has to be validated, gated on a person when it matters, idempotent, reversible, and logged. This is the application layer, not a chatbot.

The approach
  1. 01Ingestion
  2. 02Retrieval
  3. 03Reasoning
  4. 04Validation and gate
  5. 05Action and audit
Focus
Operations-heavy teams in B2B software, logistics, and financial services
Problem
Data was searchable at best. Turning a retrieved answer into a correct, safe business action still fell to a person doing it by hand.
Approach
A system that retrieves, reasons, then proposes an action, validates it, gates anything consequential on a person, and executes it once with an audit trail.

The problem

Most teams that invest in search or a RAG assistant end up in the same place: the data is now findable, and the answer arrives with its sources. Useful, and still one step short. Someone reads the answer, decides what to do, and then does it by hand: updates the record in the system, files the exception, routes the approval, kicks off the next task. The reasoning and the action still live in a person's head and hands.

The hard part was never the answer. It is the step after it. An action changes state in a real system, and a wrong one is not a bad sentence on a screen. It is a corrupted record, a duplicate charge, a decision routed to the wrong desk. Making data actionable means building the layer where a reasoned result safely does something, not just describes something.

The reality behind the problem

Wiring a model directly to the tools that write to production looks like progress and fails the moment reasoning is wrong. A model that can both decide and execute will, sooner or later, decide incorrectly and execute anyway, confidently, at machine speed, with no one between the mistake and the record. Retrieval accuracy does not save you here, because the failure is in the action, not the lookup.

The quiet failures are worse than the loud ones. The same instruction runs twice and books the action twice. A step half-completes and leaves the record in a state no one designed for. A reasonable-looking decision is taken on stale data. None of these throw an error. They just leave the business a little more wrong than it was, and no one notices until the reconciliation does.

What was assumed

If the model can find the answer, it can just do the thing.

What Stallwart asked

What exactly does it write, who approves it, and can we undo it and see who did what?

The desired outcome

Before any tool was connected to a system that writes, the engagement defined what had to be true for an action to be taken at all.

  • Every action is grounded in retrieved, current data, never in the model's general training alone.
  • An action is proposed with its reasoning attached, so a person sees why before it happens.
  • Anything consequential is held at a human gate and never executed without explicit approval.
  • The same request cannot take the same action twice, even on a retry or a duplicate trigger.
  • Every action is reversible or has a defined compensating step, so a wrong one can be undone.
  • Each action is logged with its inputs, its reasoning, its approver, and its result, queryable after the fact.

The system we designed

Stallwart builds the application layer between retrieval and the systems of record. Company data is ingested and made retrievable, as in any grounded build. The difference is what happens next: the system reasons over the retrieved records to propose a specific action, validates that action against the rules and the current state of the target system, and only then presents it. Nothing reaches a production system as a side effect of a sentence.

Consequential actions stop at a human gate. The system drafts the change, routes it with its reasoning and its citations, and waits for approval before it writes. When it does write, it writes once, idempotently, and records the full trail. Low-stakes, reversible steps can run on their own under a confidence threshold. The standard Stallwart holds is the same across both: the system never takes an action it cannot explain, cannot undo, or cannot account for.

How it works

  1. 01

    Retrieve, then reason toward an action

    The system pulls the current records and passages the decision depends on, then reasons over them to a specific proposed action: this field set to this value, this exception flagged, this task routed here. The proposal carries the reasoning and the sources it rests on, so the step is inspectable before it is real.

  2. 02

    Validate before anything is allowed to act

    The proposed action is checked against business rules and the live state of the target system before it is offered: is the record still in the expected state, does the change violate a constraint, is this a duplicate of something already done. An action that fails validation never reaches the gate.

  3. 03

    Gate, execute once, and record

    Consequential actions wait for a person's approval; reversible low-stakes ones run under a confidence threshold. Execution is idempotent, so a retry or a repeated trigger cannot double-act, and the inputs, reasoning, approver, and result are written to an audit trail.

How it works

From retrieved data to a validated action the business actually takes.

  1. 01

    Ingestion

    Every source, normalized, permissioned, current

  2. 02

    Retrieval

    The records and passages the decision rests on

  3. 03

    Reasoning

    A specific action proposed, with its rationale

  4. 04

    Validation and gate

    Checked against rules, held for approval

  5. 05

    Action and audit

    Executed once, reversible, fully logged

Each stage turns loose data into a narrower commitment, ending in one executed, logged action or an item parked for a person.

Architecture

What sits underneath.

Underneath the workflow, the system runs on the same four layers as every Stallwart build, so a guarantee made in one layer holds across the whole system.

01

Intelligence

Connectors ingest each source, normalize formats, and keep the data current and permissioned so retrieval returns records that reflect the live state, not a stale snapshot. This is the grounding the reasoning stands on, and the reason an action is never proposed on data the system has not actually checked.

02

Orchestration

Retrieval, reasoning, validation, gating, and execution run as explicit, ordered steps rather than one opaque model call. Each step has a defined input and output, so a proposed action can be inspected, held, replayed, or stopped at the point it is wrong instead of after it has written.

03

Governance

Policy decides what may run automatically and what must stop at a human gate, who can approve which class of action, and which permissions bound the data a decision may use. Every action carries its reasoning, its approver, and its sources into a trail that stays queryable long after the action is taken.

04

Production

Idempotent writes, reversibility or compensating steps, rollback, observability over every action, and an evaluation harness that grades proposed actions against known-good outcomes before a change ships. This is the layer that makes writing to a real system safe to run every day.

The hard parts

Where the engineering judgment was.

01

Separate deciding from doing

A model that both reasons and executes in one step removes the only place a mistake can be caught before it becomes a record. Splitting proposal from execution puts validation and, where it matters, a person between the reasoning and the write, which is the whole point of acting safely on reasoned output.

The trade-off

It is more moving parts than a single tool-calling agent, and it adds a step of latency. That cost buys the ability to inspect and stop an action before it happens, which on anything consequential is not optional.

02

Make every action idempotent

In production, triggers fire twice, requests retry, and queues redeliver. If an action is not idempotent, the second delivery books it again: a duplicate charge, a double update, a record the business has to reconcile by hand. Idempotency keys make the second attempt a safe no-op.

The trade-off

Every action needs a stable identity and a dedupe check, which is more design per action than fire-and-forget. It is the difference between a system that is safe to retry and one that is only safe if nothing ever goes wrong.

03

Gate on consequence, not on confidence alone

A high confidence score is not permission to act. What decides whether a person approves is the cost of being wrong and whether the action can be undone. Mapping actions by consequence, not just by model certainty, is what keeps automation aggressive where it is cheap and conservative where it is not.

The trade-off

Some reversible actions a model is sure about still wait for a person early on, which feels slower than it could be. The gate loosens as the evaluation harness earns the trust, deliberately, rather than being assumed up front.

Reliability & guardrails

How it avoids the wrong call.

The guardrails exist so that the safe outcome is the default one, and a wrong decision stops before it becomes a wrong action.

Validation before action

A proposed action is checked against business rules and the current state of the target system before it is offered. If the record has moved, a constraint would break, or the change duplicates one already made, the action is rejected or re-reasoned rather than written.

Human gate on consequence

Anything that changes state in a way that is costly or hard to reverse stops for explicit approval, routed with its reasoning and sources. The gate is a hard stop in the orchestration, not a suggestion the model can talk itself past.

Idempotent, reversible execution

Each action runs under an idempotency key so a retry or duplicate trigger cannot act twice, and every action is either reversible or paired with a defined compensating step, so a wrong one can be backed out rather than lived with.

Confidence thresholds on automation

Only reversible, low-consequence actions run without a person, and only above a confidence threshold. Below it, or outside that class, the action becomes a proposal for review instead of an execution, so thin reasoning never writes on its own.

Productionization

What turns it from a demo into software.

What separates a convincing demo from a system a team trusts to touch its records every day.

Evaluation harness for actions

A graded set of real cases runs continuously, scoring not just whether the answer was right but whether the proposed action was the correct one, so a change to the prompt, model, or rules is measured against known-good outcomes before it ships.

Audit trail by default

Every action records its inputs, the retrieved data it used, its reasoning, its approver, and its result. When an action turns out wrong, the team can see exactly which data and which step produced it, and prove who approved what.

Observability over the action path

Retrieval, reasoning, validation, gate decisions, and executions are traced end to end, so a stuck approval, a rising rejection rate, or a drift in proposed actions is visible as a signal rather than discovered in a reconciliation.

Rollback and compensation

Actions are built to be undone: a direct reversal where the target system allows it, a defined compensating step where it does not. A bad batch can be backed out as a known procedure instead of a manual cleanup.

What changed

The operational change is in who carries the last step. Instead of a person reading an answer and then doing the work by hand, the system proposes the action with its reasoning, a person approves what matters, and the write happens once, cleanly, with a record of it. The judgment stays with people; the repetitive, error-prone execution does not.

Because the system validates, gates, and logs, it earns the trust required to let it write at all. People can see what it intends to do before it does it, undo it if it was wrong, and account for every action after the fact. That is what makes automation on real systems something a team is willing to run.

  • A retrieved answer no longer stops at the screen; it becomes a proposed, validated action.
  • Consequential actions wait for a person, so automation never writes past where judgment is needed.
  • Repeated triggers and retries stop double-acting, so the system is safe to run under real conditions.
  • Every action is reversible and logged, so a wrong one can be undone and accounted for, not just regretted.

Ownership & handover

What you receive, and keep.

At handover, the system is yours to run and extend, with nothing held back.

  • Source code for the ingestion, retrieval, reasoning, and action pipeline.
  • Infrastructure as code for the stores, services, connectors, and action executors.
  • The evaluation set and harness, so you can keep grading proposed actions as rules change.
  • Runbooks for adding an action, setting its gate and confidence policy, and rolling one back.
  • Documentation of the architecture, the gating and permission model, and the audit trail.

What we learned

The lesson that generalizes: making data actionable is an application-engineering problem, not a model problem. Retrieval and reasoning are the parts a demo shows off. The value, and the risk, is in the step after: validating the action, gating it on a person where consequence demands it, making it idempotent and reversible, and recording it. A system that answers is useful. A system that acts, safely, is the one that changes how the work gets done.

This work connects to

  • AI Infrastructure & RAG
  • Retrieval and reasoning
  • Agentic action and automation
  • Human-in-the-loop systems
  • Production AI systems

Frequently asked

What does it mean to make company data actionable, beyond searchable?

Searchable means a person can find and read an answer. Actionable means the system can take the next step the answer implies: update a record, route a decision, flag an exception, trigger a downstream task. That step writes to a real system, so it has to be validated, gated on a person when it is consequential, idempotent, reversible, and logged. The answer is the easy half; the safe action is the engineering.

How is this different from a RAG assistant that answers questions?

A RAG assistant retrieves and answers, and the person decides and acts. This goes one step further: the system reasons over the retrieved data to propose a specific action, validates it against the rules and the live state of the target system, and executes it once behind a human gate. The difference is the application layer that turns a reasoned result into a committed action, not the retrieval.

How do you stop an AI from taking a wrong or irreversible action?

By separating deciding from doing. The model proposes an action; it does not execute directly. The proposal is validated against business rules and the current state before it is offered, anything consequential stops for explicit human approval, execution is idempotent so retries cannot double-act, and every action is reversible or paired with a compensating step. A wrong decision is caught as a proposal, not discovered as a corrupted record.

What is idempotency and why does it matter for automated actions?

Idempotency means running the same action twice has the same effect as running it once. In production, triggers fire twice, requests retry, and queues redeliver, so without it the system books a duplicate charge or a double update. Each action runs under a stable idempotency key and a dedupe check, which makes a repeated attempt a safe no-op and the system safe to retry.

Does a person still stay in control of what the system does?

Yes. Actions are mapped by consequence: anything costly or hard to reverse stops at a human gate and is never written without approval, routed with its reasoning and sources so the decision is informed. Only reversible, low-stakes actions run on their own, and only above a confidence threshold. Every action, automated or approved, is logged with its approver, so control and accountability stay with people.

Your data is findable. Is it doing anything, or is a person still doing it by hand?

Bring the data and the step that follows it. We build the layer that reasons, proposes, and takes the action, validated, gated, and logged.

Last updated: October 6, 2026