Skip to content

Case study · AI Infrastructure & RAG

Turning scattered knowledge into an assistant that cites its sources

The knowledge already existed. It was spread across PDFs, wikis, shared drives, and a few internal tools, so finding an answer meant knowing where to look and who to ask. The problem was never a missing document. It was retrieval, grounding, and knowing when to say it was not sure.

The approach
  1. 01Ingestion
  2. 02Chunking
  3. 03Retrieval and rerank
  4. 04Grounded generation
  5. 05Citations and confidence
Focus
Knowledge-heavy operations in professional services and B2B software
Problem
Answers lived across PDFs, docs, wikis, and internal tools, so finding one meant knowing where to look.
Approach
A retrieval system that grounds every answer in source passages, cites them, and escalates when confidence is low.

The problem

Most of the knowledge needed to answer day-to-day questions already existed inside the business. It just lived in too many places: product PDFs, a documentation site, a wiki, shared drives, support history, and a couple of internal tools that each held part of the picture. Answering a question meant knowing which source held it, and often which person had read it last.

A plain chatbot bolted onto one of those sources does not fix this. The hard part is retrieving the right passage from the right source at the right freshness, grounding the answer in it, and being honest when the sources disagree or simply do not hold the answer.

The reality behind the problem

The obvious fix, point a language model at the documents, fails in the ways that matter in production. Raw documents are not retrievable units: a forty-page PDF answers thirty different questions, and dropping the whole file into a prompt buries the one relevant paragraph. Sources contradict each other, a deprecated policy still sits in an old file next to the current one, and nothing tells the model which to trust.

The failure mode is quiet. A retrieval step that returns nothing useful does not make the system go silent; the model fills the gap with a fluent, plausible, wrong answer. Without grounding and a way to measure its own confidence, the system is most dangerous exactly when it knows the least.

What was assumed

We need a chatbot on our documents.

What Stallwart asked

Which passage answers this, from which source, and how sure are we?

The desired outcome

Before a model was chosen, the engagement defined what the system actually had to do to be trusted in daily use.

  • Answer from the company's own sources, never from the model's general training alone.
  • Show the passages and documents an answer came from, so a person can verify it.
  • Prefer the current source when two disagree, and surface the conflict rather than hide it.
  • Say it is not sure and route to a person when retrieval is weak, instead of guessing.
  • Stay current as documents change, without a full re-index by hand each time.
  • Respect who is allowed to see what, so the assistant never answers past a permission boundary.

The system we designed

Stallwart builds a retrieval-augmented system that treats company knowledge as a governed source of truth, not a prompt to stuff. Documents from every source are ingested, normalized, and split into passages sized to answer a question, each tagged with its source, its last-updated date, and the permissions that govern it. At query time the system retrieves the passages that actually match, grounds the answer in them, and returns the citations alongside it.

The assistant is built to report its own uncertainty. When retrieval is thin or the top passages disagree, it says so and escalates to a person rather than inventing a resolution. That is the standard Stallwart holds across every build: a system that surfaces what it does not know, and keeps an auditable record of what it answered and from where.

How it works

  1. 01

    Make documents retrievable

    Every source is ingested, normalized, and split into passages sized to answer a question, each carrying its source, freshness, and access scope. Retrieval quality is decided here, before the model ever runs.

  2. 02

    Retrieve, then rerank

    A query pulls candidate passages from the index, then a reranking pass orders them by whether they actually answer the question, not just whether they look similar. Permission filters drop anything the asker cannot see.

  3. 03

    Ground, cite, and gate

    The model answers strictly from the retrieved passages and returns its citations. A confidence signal gates the response: below the threshold, the assistant escalates to a person instead of guessing.

How it works

From scattered sources to a grounded, cited answer.

  1. 01

    Ingestion

    Every source, normalized and permissioned

  2. 02

    Chunking

    Passages sized to answer a question

  3. 03

    Retrieval and rerank

    The few passages that truly match

  4. 04

    Grounded generation

    Answer written only from retrieved text

  5. 05

    Citations and confidence

    Sources shown, low confidence escalated

Each stage narrows an open question to a specific, attributable passage, or to an honest 'I am not sure.'

Architecture

What sits underneath.

Underneath the workflow, the assistant runs on the same four layers as every Stallwart build, so a guarantee made in one layer holds across the whole system.

01

Ingestion and indexing

Connectors pull from each source on a schedule, normalize formats such as PDF, HTML, docs, and tickets, split content into passages, and write embeddings and metadata, source, timestamp, and access scope, into a vector index and a document store.

02

Retrieval and reranking

A query is embedded and matched against the index, a reranking pass reorders candidates by true relevance rather than raw similarity, and permission filters drop anything the asker is not allowed to see before a single token is generated.

03

Grounded generation

The model answers from the retrieved passages only, under instructions that forbid filling gaps from general knowledge, and returns the citations it used so every claim traces back to a source a person can open.

04

Governance and observability

Every query, the passages it retrieved, the answer, and the confidence signal are logged. Low-confidence or conflicting results route to a person, and the whole trail stays queryable after the fact.

The hard parts

Where the engineering judgment was.

01

Chunking is most of the engineering

Retrieval quality is set long before the model runs. Passages that are too large bury the answer; too small, and they lose the context that makes them meaningful. Getting chunk boundaries right on real documents is the bulk of the work.

The trade-off

There is no universal chunk size. It is tuned per source type and checked against real questions, which takes iteration rather than a default.

02

Rerank, do not trust raw similarity

Vector similarity surfaces passages that look related but do not answer the question. A reranking pass is what separates a plausible retrieval from a correct one.

The trade-off

An extra model call per query adds latency and cost, spent on purpose because retrieval accuracy is what the whole system stands on.

03

Refuse before you guess

The most valuable behavior is the assistant declining when it cannot ground an answer. One wrong answer that looks right erodes trust in every answer after it.

The trade-off

The system will sometimes say it is not sure on a question a person could have answered, a deliberate bias toward silence over confident error.

Reliability & guardrails

How it avoids the wrong call.

The guardrails exist to make the honest failure mode the default one.

Grounding, enforced

Answers are constrained to the retrieved passages. If the sources do not contain the answer, the system does not improvise one from the model's general training.

Confidence thresholds

A retrieval and agreement score gates each answer. Below the threshold, the assistant escalates to a person instead of responding, so thin information never reads as high confidence.

Conflict surfacing

When two current sources disagree, the assistant shows both and flags the conflict rather than quietly picking one, turning a hidden contradiction into a visible decision.

Permission-aware retrieval

Access scope is applied at retrieval, not after generation, so the assistant cannot compose an answer from a document the asker is not allowed to read.

Productionization

What turns it from a demo into software.

What separates a convincing demo from a system a team relies on every day.

Evaluation harness

A graded set of real questions runs continuously, so a change to the prompt, model, or index is measured against known-good answers before it ships, not discovered in production.

Incremental re-indexing

When a document changes, only its passages are re-embedded, so the knowledge base stays current without a full manual rebuild and answers do not drift from the live sources.

Observability

Each answer carries its retrieved passages, citations, and confidence into the logs, so when an answer is wrong the team can see which passage and which step produced it.

Cost and latency control

Retrieval depth, reranking, and model choice are tuned to hold answer latency and per-query cost inside a known budget rather than letting them float.

What changed

The operational change is in where the knowledge lives and how it is reached. Instead of an answer depending on who had read which document, a question returns a grounded answer with its sources attached, and a verifiable 'I am not sure' when the sources do not hold it.

Because the system cites and escalates, it earns the kind of trust a confident black box never does. People can check it, and it tells them when not to rely on it.

  • Finding an answer no longer depends on knowing which document or person holds it.
  • Every answer arrives with its sources, so it can be verified rather than trusted blind.
  • Low-confidence questions route to a person instead of producing a fluent wrong answer.
  • The knowledge base stays current as documents change, with no manual rebuild.

Ownership & handover

What you receive, and keep.

At handover, the system is yours to run and extend, with nothing held back.

  • Source code for the ingestion, retrieval, and generation pipeline.
  • Infrastructure as code for the vector index, document store, and services.
  • The evaluation set and harness, so you can keep grading answers as sources grow.
  • Runbooks for adding a source, re-indexing, and tuning retrieval.
  • Documentation of the architecture, the guardrails, and the permission model.

What we learned

The lesson that generalizes: a knowledge assistant is a retrieval and governance problem long before it is a model problem. The model is the least differentiated part. The value is in how documents become retrievable, how confidence is measured, and how honestly the system declines, which is exactly the work a demo skips and production demands.

This work connects to

  • AI Infrastructure & RAG
  • Retrieval-augmented generation
  • Grounding and citations
  • Knowledge systems
  • Production AI systems

Frequently asked

What is a RAG knowledge assistant?

A RAG (retrieval-augmented generation) knowledge assistant answers questions by first retrieving relevant passages from a company's own documents and data, then generating an answer grounded in those passages and citing them. It answers from your sources rather than from the model's general training alone, which is what makes its answers verifiable.

How does a RAG assistant avoid hallucinating?

By grounding every answer in retrieved source passages, citing them, and measuring its own confidence. When retrieval is weak or sources conflict, a well-built assistant says it is not sure and escalates to a person instead of inventing an answer. Grounding plus a confidence gate is what turns a fluent guess into a checkable answer.

How is this different from putting a chatbot on our documents?

A chatbot on one source can only answer from that source and usually has no way to cite, to resolve conflicts, or to know when it is wrong. A production RAG system ingests every source, retrieves and reranks the passages that actually match, enforces permissions, grounds and cites the answer, and escalates low-confidence questions. The difference is retrieval quality and governance, not the model.

How does the assistant stay current as our documents change?

Sources are re-ingested on a schedule, and when a document changes only its passages are re-embedded, so the knowledge base tracks the live sources without a full manual rebuild. That keeps answers aligned with the current version rather than a stale snapshot.

Can it respect who is allowed to see what?

Yes. Each passage carries the permissions of its source, and access scope is applied at retrieval, before an answer is generated. The assistant can only compose an answer from documents the person asking is allowed to see, so it never answers past a permission boundary.

Is the answer already somewhere in your company, just not where anyone can find it?

Bring the sources. We build a system that retrieves the right one, cites it, and tells you when it is not sure.

Last updated: October 6, 2026