Skip to content

Case study · AI Agents & Automation

The lead-research system that works before the salesperson does

Every outbound touch started with a person doing research by hand: finding the account, pulling the firmographics, checking for a reason to reach out, then writing something specific. That research gated every conversation, and it did not scale. The work was real, repeatable, and almost entirely mechanical, right up to the point where judgment actually mattered.

The approach
  1. 01Discovery
  2. 02Enrichment
  3. 03Qualification
  4. 04Personalization
  5. 05Handoff
Focus
Outbound B2B sales and revenue operations
Problem
Research had to happen by hand before every outbound touch, so it gated each conversation and could not scale.
Approach
An agent pipeline that discovers accounts, enriches them through external APIs, scores fit and intent, and hands a context-rich lead to a person.

The problem

Before a salesperson could send a single outbound message, someone had to do the research. Find the right accounts, resolve the right contacts, pull firmographic detail, look for a recent signal worth mentioning, and only then write something that did not read like a template. That preparation was the bottleneck. It happened one account at a time, and it happened before every conversation, so the number of conversations a team could start was capped by how fast people could research.

The work itself was mechanical. Most of it was finding, resolving, and cross-checking facts that already existed in external sources. The judgment, deciding whether an account is worth a human's time and what to open with, came at the very end, after hours of lookup. So the scarce resource, a salesperson's attention, was spent mostly on data gathering rather than on the conversation it was meant to enable.

The reality behind the problem

The obvious fix, hand the whole thing to one model and ask it to research a list of companies, breaks on contact with real data. A language model on its own does not know today's headcount, the current funding round, or who holds a role this quarter. Asked anyway, it produces fluent, specific, confidently wrong facts, and a wrong fact in an outbound message is worse than a generic one because it signals the sender never checked.

Pulling live data is not enough on its own either. External sources disagree, go stale, and describe the same company under three different names. The same account arrives as a duplicate from two providers. A firmographic field is six months out of date. Without a layer that resolves entities, reconciles conflicting sources, and refuses to assert a fact it cannot support, automation just produces bad research faster than a person ever could.

What was assumed

We need a model that writes our outbound messages.

What Stallwart asked

Is this account worth a person's time, what do we actually know about it, and how current is that?

The desired outcome

Before any model or data provider was chosen, the engagement defined what the pipeline had to do to be trusted ahead of a live conversation.

  • Resolve each target to a real, de-duplicated account and the right current contacts, not a near-match or a stale record.
  • Enrich only from sources the system can attribute, and never assert a prospect fact it cannot trace back to one.
  • Score fit and intent explicitly, so a person sees why an account surfaced and can disagree with it.
  • Ground every personalized message in a fact the research actually found, with that fact on the record.
  • Hand off a qualified, context-rich lead to a person, with the conversation itself left to the person.
  • Keep an auditable trail of what was pulled, from where, when, and how the lead was scored.

The system we designed

Stallwart builds the pre-conversation pipeline end to end. Agents discover and resolve target accounts and their current contacts, enrich each one through external firmographic and signal APIs, score the account for fit and intent, and draft personalization grounded in what the research actually returned. Each stage writes what it found and where it found it, so a lead arrives not as a name on a list but as a resolved account with attributed facts, an explicit score, and an opener tied to a real signal.

The pipeline is built to know the limits of its own data. Facts are only asserted when a trusted source supports them, conflicting sources are reconciled rather than averaged, and an account the system cannot confidently research or score is routed to a person instead of pushed forward on a guess. The system does the research that used to gate the conversation. A person still runs the conversation.

How it works

  1. 01

    Discover and resolve, do not just list

    Agents turn a target definition into real accounts and contacts, then resolve each to a single canonical record: de-duplicating providers, matching entities across naming variants, and confirming a contact's current role. Resolution quality is decided here, before any enrichment or scoring runs on top of it.

  2. 02

    Enrich from sources you can attribute

    Each resolved account is enriched through external firmographic and signal APIs. Every field carries the source it came from and when it was pulled. Where sources conflict, the pipeline reconciles toward the most current trusted one and records the conflict rather than silently choosing.

  3. 03

    Score, ground, and hand off

    The account is scored for fit and intent with the contributing signals attached, so a person sees why it surfaced. Personalization is drafted strictly from the facts the research found. A qualified, context-rich lead is routed to a person; a low-confidence account is escalated instead of sent.

How it works

From a target account to a lead a person can open a conversation with.

  1. 01

    Discovery

    Target accounts and contacts found and resolved

  2. 02

    Enrichment

    Firmographic and signal data pulled via external APIs

  3. 03

    Qualification

    Fit and intent scored, with the reasons recorded

  4. 04

    Personalization

    Each message grounded in a fact the research found

  5. 05

    Handoff

    A context-rich lead routed to a person

Each stage turns an unverified target into a resolved, scored, grounded lead, or routes it to a person when the research will not support a confident answer.

Architecture

What sits underneath.

Underneath the workflow, the pipeline runs on the same four layers as every Stallwart build, so a guarantee made in one layer holds across the whole system.

01

Intelligence

Agents handle discovery, entity resolution, and the reading of enrichment data: matching an account across naming variants, reconciling records from more than one provider, and judging which returned facts are usable rather than treating every API response as ground truth.

02

Orchestration

A pipeline sequences discovery, enrichment through external APIs, qualification scoring, and personalization as discrete stages, with retries, rate-limit handling, and caching so a flaky or slow data source degrades one account rather than stalling the run.

03

Governance

Every field carries its source and timestamp, every score carries the signals behind it, and conflicting or stale data is flagged. An account the system cannot support is routed to a person, and the full trail stays queryable after the fact.

04

Production

The pipeline runs against real provider quotas and real data drift: scheduled refreshes, an evaluation set that checks resolution and scoring against known-good accounts, observability over each stage, and the ability to roll a change back cleanly.

The hard parts

Where the engineering judgment was.

01

Entity resolution is most of the engineering

A lead pipeline stands on knowing that two records are the same company and that a contact still holds the role. Get resolution wrong and every later stage enriches, scores, and personalizes the wrong account with perfect confidence. The bulk of the work is in matching and de-duplication, not in writing the message.

The trade-off

There is no universal match rule. Resolution is tuned per source and checked against real accounts, which takes iteration rather than a default, and some genuine matches are held for review rather than merged blind.

02

Attribute every fact or do not assert it

Personalization is only worth sending if the fact behind it is real. Binding each enriched field to a source and timestamp is what lets the system refuse to state something it cannot support, instead of letting a model fill the gap with a plausible guess.

The trade-off

The pipeline will sometimes return a thinner profile than a provider seems to offer, because an unattributable or stale field is dropped rather than used. Less surface, more trust.

03

Score in the open, and hand off to a person

Fit and intent are judgments, and a judgment a salesperson cannot see is one they cannot correct. Exposing the signals behind each score, and routing the lead to a person for the actual conversation, keeps the human in the loop where judgment belongs.

The trade-off

An explainable score is not always the highest-performing one in the abstract, and some accounts a model might have pushed are instead escalated for a person to decide. The system optimizes for a lead a person can trust over raw throughput.

Reliability & guardrails

How it avoids the wrong call.

The guardrails exist to make the honest failure mode the default one: a thin, flagged lead rather than a confident wrong one.

Source trust and attribution

Every enriched fact is bound to the provider it came from and the time it was pulled. The pipeline asserts a prospect fact only when a trusted source supports it, so an answer can always be traced back rather than taken on faith.

Deduplication and entity resolution

Records from different sources are matched and merged into one canonical account before anything is scored or sent, so the same company is not researched twice and a contact is tied to the right organization.

Stale-data handling

Fields carry a freshness signal, and data past its useful age is refreshed or dropped rather than used. When two sources disagree, the pipeline reconciles toward the current one and records the conflict instead of hiding it.

No fabricated prospect facts

The model drafts personalization strictly from facts the research returned. If the research did not find a usable signal, the system does not invent one, and the account is handed off plain or escalated rather than dressed up with a guess.

Productionization

What turns it from a demo into software.

What separates a convincing demo from a pipeline a sales team runs every day.

Evaluation harness

A graded set of known accounts checks entity resolution, enrichment accuracy, and scoring against expected answers, so a change to a prompt, a provider, or a scoring rule is measured before it ships rather than discovered in sent outbound.

API quota and rate-limit control

Calls to external data providers are batched, cached, and rate-limited to a known budget, so a run stays inside its provider quotas and a single slow source degrades gracefully instead of failing the whole pipeline.

Observability

Each lead carries its sources, the data pulled, the score and its signals, and the drafted message into the logs, so when a lead is wrong the team can see which stage, which source, and which fact produced it.

Scheduled refresh and rollback

Accounts are re-enriched on a schedule so leads track the live data, and any change to the pipeline can be rolled back cleanly if a provider shifts its schema or a scoring change misfires.

What changed

The operational change is in where research sits relative to the conversation. Instead of a person spending most of their time finding and verifying facts before they can reach out, the pipeline arrives at a resolved, scored, grounded lead first, and the person starts from there. Research stops being the gate on how many conversations can begin.

Because every fact is attributed and every score is explained, a salesperson can trust what they are handed and can see when not to. The attention that used to go to lookup goes to the conversation, and the conversation stays a human one.

  • Research runs ahead of the salesperson instead of gating each outbound touch.
  • Every lead arrives with attributed facts and an explained score, so it can be checked rather than trusted blind.
  • Accounts the system cannot confidently research or score route to a person instead of going out on a guess.
  • A salesperson's time shifts from gathering data to the conversation the data was meant to start.

Ownership & handover

What you receive, and keep.

At handover, the pipeline is yours to run and extend, with nothing held back.

  • Source code for the discovery, enrichment, qualification, personalization, and handoff pipeline.
  • Infrastructure as code for the services, data store, and external API integrations.
  • The evaluation set and harness, so you can keep grading resolution and scoring as sources change.
  • Runbooks for adding a data provider, tuning the scoring model, and handling a source outage.
  • Documentation of the architecture, the attribution and deduplication model, and the handoff rules.

What we learned

The lesson that generalizes: an outbound lead-research system is a data-resolution and governance problem long before it is a writing problem. The message is the least differentiated part. The value is in resolving an account to a single trustworthy record, attributing every fact, scoring in the open, and refusing to assert what the sources do not support, which is exactly the work a demo skips and a sales team feels every day.

This work connects to

  • AI Agents & Automation
  • Data enrichment pipelines
  • Outbound sales automation
  • Human-in-the-loop systems
  • Production AI systems

Frequently asked

What does an AI lead-research system actually do?

It runs the research that happens before an outbound conversation: finding and resolving target accounts and contacts, enriching them with firmographic and signal data from external APIs, scoring fit and intent, and drafting personalization grounded in what the research found. It hands a context-rich lead to a salesperson, who runs the actual conversation. The system does the preparation that used to gate each outreach, not the selling.

How is this different from a tool that writes outbound emails?

An email writer starts from whatever facts it is given and produces copy. A lead-research pipeline produces the facts first: it resolves the real account, enriches it from attributable sources, scores it, and only then grounds a message in a signal it actually found. The message is the last and least differentiated step. The work and the value are in the research and the governance underneath it.

How does it avoid sending outbound built on wrong or made-up facts?

Every enriched field is bound to the source it came from and when it was pulled, and the pipeline asserts a prospect fact only when a trusted source supports it. Personalization is drafted strictly from facts the research returned, so if no usable signal was found the system does not invent one. The account is handed off plain or escalated instead of dressed up with a guess.

Does a person still run the conversation?

Yes. The system does the pre-conversation research and hands a qualified, context-rich lead to a person, who takes the conversation from there. Accounts the pipeline cannot confidently research or score are routed to a person rather than pushed forward automatically. The human stays in the loop for the judgment and the conversation itself.

How does it handle duplicate and stale data from different providers?

Records from different sources are matched and merged into one canonical account before anything is scored or sent, so the same company is not researched twice. Fields carry a freshness signal, and data past its useful age is refreshed or dropped. When two sources disagree, the pipeline reconciles toward the current trusted one and records the conflict rather than hiding it.

Is your team's outreach capped by how fast people can research?

Bring the accounts. We build a pipeline that resolves them, enriches them from sources it can attribute, scores them in the open, and hands your team a lead worth a conversation.

Last updated: October 6, 2026