Skip to content

Case study · AI Agents & Automation

The read-and-key work, done by a system that validates every field

Contracts, invoices, forms, reports, and applications arrive in every layout and every scan quality. Someone reads each one, keys the fields into a system, checks them, and handles the odd ones. The work is not hard. It is slow, repetitive, and exactly where a wrong value does the most damage once it is written downstream.

The approach
  1. 01Intake and classify
  2. 02Extraction
  3. 03Validation
  4. 04Confidence and exceptions
  5. 05Downstream write
Focus
Document-heavy back-office operations across finance, insurance, and logistics
Problem
People read each document and key its fields by hand, then check them, then handle the odd ones.
Approach
A system that extracts the fields, validates each one against rules and sources, and routes low-confidence cases to a person before anything is written.

The problem

The work behind document-heavy operations is the same loop repeated all day. A document arrives, a person reads it, finds the fields that matter, keys them into a system, checks them against what they already know, and files the ones that do not fit for someone to look at later. Across contracts, invoices, forms, reports, and applications, this is where a large share of operational time goes, and none of it is the kind of work a person is good at for long.

A simple extractor pointed at one clean template does not fix this. Real documents do not arrive clean. They come in dozens of layouts, from different senders, as photos, faxes, and scans of varying quality, with handwriting, stamps, and fields in the wrong place. The hard part is reading them reliably, then proving each extracted value is right before it moves on.

The reality behind the problem

The obvious fix, run optical character recognition and map the text to fields, holds up on the sample documents and breaks on the real ones. Layout varies from sender to sender, so a rule that finds the total on one invoice misses it on the next. Scan quality degrades the text before any model sees it. A field read with 60 percent certainty and a field read with 99 percent certainty look identical once they land in a cell, and the system has no way to tell them apart.

The dangerous failure is silent. An extractor that misreads a date or a number does not stop. It writes a confident, plausible, wrong value straight into the system that consumes it, and the error surfaces weeks later in a payment, a contract term, or an application decision. Without validation and a measured confidence on every field, the system is least trustworthy exactly where it looks most finished.

What was assumed

We need an extractor that reads our documents.

What Stallwart asked

For each field, what did we read, how sure are we, and does it pass the checks before it is written?

The desired outcome

Before any model was chosen, the engagement defined what the system actually had to do to replace the manual read-and-key step with something a team could trust.

  • Extract the required fields from documents across varied layouts, senders, and scan quality, not one clean template.
  • Validate every extracted field against format rules, business logic, and the systems of record, not just read it.
  • Attach a confidence to each field and each document, so a weak read is never mistaken for a strong one.
  • Auto-process only what clears the threshold, and route everything below it to a person with the document and the flagged fields in view.
  • Never write an unvalidated or low-confidence value into a downstream system silently.
  • Keep an auditable record of every field: what was read, from where on the page, how it validated, and who, if anyone, reviewed it.

The system we designed

Stallwart builds a document-processing system that does the reading and the keying, and treats validation as the point rather than an afterthought. Documents are taken in from every channel, classified by type, and read field by field, with each value carrying where on the page it came from and how confident the extraction is. Every field is then checked against format rules, cross-field logic, and the systems that already hold related data, so a value is only trusted once it has passed.

The system is built to know what it is unsure of. Fields and documents below the confidence threshold, and any that fail validation, are routed to a person with the original document and the flagged values side by side, not pushed through. Only validated data is written downstream, and every decision is logged. That is the standard Stallwart holds across every build: a system that surfaces what it cannot stand behind, keeps a person in the loop where judgment is needed, and leaves an auditable trail of what it did.

How it works

  1. 01

    Read the document, field by field

    Documents are classified by type, then read for the fields each type requires, across varied layouts and scan quality. Every extracted value carries where on the page it came from and how confident the read is, so the system can reason about its own output instead of treating all values alike.

  2. 02

    Validate before trusting

    Each field is checked against format rules, cross-field logic, and the systems of record that already hold related data. A total that does not sum, a date outside a valid range, a supplier that does not match the register: each is caught here, before the value is allowed to move.

  3. 03

    Route by confidence, then write

    Fields and documents that clear the confidence threshold and pass validation are written downstream automatically. Everything else is routed to a person with the original document and the flagged fields in view. Nothing unvalidated is written silently.

How it works

From a document of any layout to a validated, attributable record.

  1. 01

    Intake and classify

    Every channel, every layout, sorted by document type

  2. 02

    Extraction

    Fields read with position and a per-field confidence

  3. 03

    Validation

    Checked against rules, cross-field logic, and systems of record

  4. 04

    Confidence and exceptions

    Low-confidence or failing fields routed to a person

  5. 05

    Downstream write

    Only validated data pushed into the systems that consume it

Each stage turns a page into fields, checks them, and either clears them to write or routes them to a person, never a quiet guess.

Architecture

What sits underneath.

Underneath the workflow, the system runs on the same four layers as every Stallwart build, so a guarantee made in one layer holds across the whole system.

01

Extraction and intelligence

Documents are taken in from each channel, classified by type, and read field by field across varied layouts and scan quality. Each extracted value is tagged with its location on the page and a confidence score, so weak reads are visible rather than hidden among strong ones.

02

Validation and orchestration

Every field is checked against format rules, cross-field logic, and the systems of record before it is trusted. The orchestration layer sequences classify, extract, validate, and route, retries what is recoverable, and holds each document's state so nothing is lost or double-processed.

03

Exception handling and human review

Fields below the confidence threshold or failing validation are routed into a review queue with the original document and the flagged values side by side. A reviewer corrects or confirms, the correction feeds back, and only then does the document continue.

04

Governance and observability

Every field carries an audit trail: what was read, from where on the page, how it validated, its confidence, and who reviewed it. Each document's full path is logged and stays queryable, so any downstream value can be traced back to its source.

The hard parts

Where the engineering judgment was.

01

Layout variation and poor scans are the real work

The documents that break a system are the ones a demo never shows: a new sender's layout, a photographed page, a faxed form with a stamp over a field. Reading these reliably, and knowing when a scan is too degraded to trust, is where most of the engineering goes.

The trade-off

There is no single extractor that handles every layout out of the box. It is tuned per document type and tested against real, messy samples, which takes iteration rather than a default.

02

Validation is the point, not extraction

Pulling a value off a page is the easy half. Proving it is right, that the total sums, the date is plausible, the party matches the register, is what lets the value be trusted without a person re-checking it. A system that extracts but does not validate just moves the checking, it does not remove it.

The trade-off

Validation rules have to be built and maintained per document type and per downstream system, which is ongoing work rather than a one-time setup.

03

Confidence thresholds decide auto versus review

The threshold is the dial that balances automation against risk. Set it to auto-process everything and wrong values reach production. Set it to review everything and nothing is saved. Tuning it per field and per document type, against the cost of each kind of error, is how the system stays both useful and safe.

The trade-off

A higher threshold sends more documents to human review and automates less. That is a deliberate bias toward routing a borderline case to a person over writing a wrong value downstream.

Reliability & guardrails

How it avoids the wrong call.

The guardrails exist to make the safe failure, routing to a person, the default one.

Validation, enforced

No field is trusted on extraction alone. It must pass format rules, cross-field logic, and a check against the systems of record. A value that cannot be validated is treated as an exception, not written as if it were confirmed.

Confidence thresholds

Each field and each document carries a confidence score, and a threshold gates what is auto-processed. Below it, the case is routed to review rather than written, so a weak read never passes as a strong one.

Human review loop

Exceptions land in a queue with the original document and the flagged fields in view, so a reviewer can correct or confirm quickly. The correction is captured, both to complete the document and to improve the extraction over time.

No silent downstream writes

The system never writes an unvalidated or low-confidence value into a consuming system. A field either clears validation and the threshold, or it waits for a person. The dangerous quiet failure is designed out.

Productionization

What turns it from a demo into software.

What separates a convincing demo from a system a team runs every day.

Evaluation harness

A labelled set of real documents, including the hard layouts and poor scans, runs continuously, so a change to extraction, validation, or the threshold is measured against known-correct fields before it ships rather than discovered in production.

New layouts and drift

When a new sender or a changed form appears, it shows up as a drop in confidence and a rise in exceptions rather than as silent errors. New document types are added to the classifier and the validation rules without rebuilding the system.

Observability

Each document carries its extracted fields, confidences, validation results, and routing decision into the logs, so when a value is wrong the team can see which field, which check, and which step produced it.

Safe writes and rollback

Downstream writes are reversible and traced, so a batch that turns out to be wrong can be identified and rolled back by document and by field rather than hunted for across the systems it touched.

What changed

The operational change is that the manual read-and-key step is handled by the system for everything that clears validation and the confidence threshold. People stop keying and checking routine documents, and spend their time on the exceptions the system deliberately routes to them, the cases that genuinely need judgment.

Because every field is validated and every decision is logged, the data written downstream can be trusted and traced rather than taken on faith. When something is wrong, it is visible and attributable, not buried in a cell that looked processed.

  • Routine documents are read, validated, and written without a person keying them by hand.
  • People work the exceptions the system flags, instead of processing every document to find them.
  • Every extracted field is validated against rules and sources before it is trusted.
  • Each value written downstream traces back to the document, the position, and the check that cleared it.

Ownership & handover

What you receive, and keep.

At handover, the system is yours to run and extend, with nothing held back.

  • Source code for the intake, classification, extraction, validation, and downstream-write pipeline.
  • Infrastructure as code for the services, queues, and stores the system runs on.
  • The evaluation set and harness, so you can keep grading extraction and validation as documents change.
  • The validation rules and confidence thresholds, documented and editable per document type.
  • Runbooks for adding a document type, tuning a threshold, and handling the review queue.

What we learned

The lesson that generalizes: automating document-heavy work is a validation and exception-handling problem long before it is an extraction problem. Reading a field off a page is the part a demo shows. Proving the field is right, measuring how sure the system is, and routing the uncertain cases to a person without ever writing a wrong value silently is the part production depends on, and the part that earns the trust to let the system run.

This work connects to

  • AI Agents & Automation
  • Intelligent document processing
  • Extraction and validation
  • Human-in-the-loop automation
  • Production AI systems

Frequently asked

What is intelligent document processing?

Intelligent document processing is the automation of the manual work behind documents like contracts, invoices, forms, and applications: reading each document, extracting the fields that matter, validating them, and writing them into the systems that consume them. A production system does all four, not just the reading, and keeps a person in the loop for the cases it cannot stand behind.

How does the system handle messy scans and varied layouts?

Documents are classified by type, then read field by field across the layouts and scan quality that actually arrive, including photos, faxes, and degraded scans. Each extracted value carries a confidence score, so a field read from a poor scan is flagged as uncertain rather than treated the same as a clean one, and uncertain cases are routed to a person.

How is validation different from extraction?

Extraction pulls a value off the page. Validation proves it is right by checking it against format rules, cross-field logic, and the systems of record that already hold related data. A system that only extracts moves the checking onto a person; a system that validates lets a field be trusted without being re-checked, which is what actually removes the manual work.

What stops a wrong value from being written into our systems?

Two things: validation and a confidence threshold. A field must pass its checks and clear the threshold before it is written. Anything that fails validation or reads with low confidence is routed to a person with the document in view, never written silently. The system is built so the safe failure, waiting for a person, is the default.

How are low-confidence or unusual documents handled?

They become exceptions. Fields below the confidence threshold or failing validation land in a review queue with the original document and the flagged values side by side, so a reviewer can correct or confirm quickly. The correction completes the document and feeds back to improve extraction, so the system handles more over time without losing the human check where it matters.

Is your team still reading documents and keying the fields by hand?

Bring the documents, in whatever shape they arrive. We build a system that reads them, validates every field, routes the uncertain ones to a person, and writes only what it can stand behind.

Last updated: October 6, 2026