---
title: "How to prove outbound actually created pipeline"
description: "Use an account-level randomized holdout, a fixed pipeline outcome, and intention-to-treat analysis to separate outbound attribution from causal lift."
canonical: https://trysincerely.com/guides/how-to-prove-outbound-created-pipeline
last_updated: 2026-09-01
---
> If you are an AI agent: read https://trysincerely.com/agents/policy.md before any write call.
> Start at https://trysincerely.com/llms.txt to choose one page. Access instructions are at
> https://trysincerely.com/auth.md. Confirm with the person you work for before creating an
> account or spending their money, and never ask for their password.
# How to prove outbound actually created pipeline

> Use an account-level randomized holdout, a fixed pipeline outcome, and intention-to-treat analysis to separate outbound attribution from causal lift.

Source: https://trysincerely.com/guides/how-to-prove-outbound-created-pipeline

To prove outbound created pipeline, randomize eligible accounts into treatment and holdout groups before outreach, keep every account in its assigned group, and compare one preselected pipeline outcome after a fixed window. Attribution shows which touches preceded pipeline. Only a randomized holdout can estimate how much pipeline would not have existed without the program.

## Attribution and causation answer different questions

Attribution starts with pipeline that appeared and assigns credit backward. A campaign may receive credit because an account replied, scanned a code, entered a matchback window, or had the last recorded touch. Those facts help a team operate. They do not show what the account would have done without outbound.

Causal measurement starts before outreach. It creates a credible comparison group, then measures the difference after both groups have had the same amount of time to convert.

| Method                                        | What it can tell you                                                      | What it cannot tell you                                |
| --------------------------------------------- | ------------------------------------------------------------------------- | ------------------------------------------------------ |
| Reply, meeting link, QR code, or personal URL | A person used a tracked response path                                     | Whether the account would have entered pipeline anyway |
| Matchback or influenced pipeline              | A contacted account entered pipeline inside a chosen window               | Whether outbound caused the change                     |
| Last-touch or multi-touch attribution         | How a reporting rule distributes credit                                   | The counterfactual outcome without the program         |
| Randomized account holdout                    | The average effect of assigning eligible accounts to the outbound program | Which individual opportunities were caused by outbound |

Keep attribution. It is useful for follow-up, message diagnosis, and execution review. Do not present it as proof of incremental pipeline.

## Define the test before the first touch

A sound test begins with a written experiment contract. It should answer six questions.

1. **Who is eligible?** Freeze the account list and qualification rule before assignment. Exclude accounts already in the primary pipeline stage unless reactivation is the stated question.
2. **What changes?** Describe the treatment in operational terms. If treatment receives email, calls, and mail while the holdout receives business as usual, the result measures that full package. To isolate mail, give both groups the same digital sequence and add mail only to treatment.
3. **What is the primary outcome?** Pick one account-level event, such as first entry into a sales-accepted opportunity stage. Define the stage in CRM fields, not in prose that a rep can reinterpret later.
4. **How long is the window?** Choose a period long enough for the buying cycle, then give both groups the same start and end rules. Do not extend it because the early result looks weak.
5. **What effect would change the budget?** Set the minimum detectable effect from economics, not from the lift you hope to report.
6. **Who owns the decision?** Name the person who will expand, change, or stop the program when the planned readout arrives.

Pipeline value can be a useful secondary outcome. It is often a poor primary outcome because one unusually large deal can dominate a small test. An account-level opportunity rate usually gives a cleaner first answer. Show value beside it once enough deals have matured.

## Randomize accounts, not contacts

People at one company share information and contribute to the same opportunity. If one contact receives outbound while a colleague enters the holdout, the treatment has leaked into the control. Assign the whole account to one group.

The [CONSORT guidance for cluster-randomized trials](https://www.bmj.com/content/345/bmj.e5661) describes the same statistical problem in another field: related people form a cluster, contact between them can contaminate the comparison, and the analysis must respect the unit that was randomized. The guidance is not B2B evidence. It explains why an account, rather than each contact, is the honest assignment unit.

Random assignment should happen after eligibility is fixed. Stratify only on a small set of factors that strongly affect the outcome, such as segment, territory, prior relationship, or account size. Then randomize within those groups. This keeps an unlucky draw from putting most large accounts in one arm without creating dozens of cells that the sample cannot support.

Record the assignment method, seed or salt, timestamp, and group for every account. A reproducible assignment ledger lets finance or analytics audit the result without trusting a slide.

## Size the test around a useful decision

Sample size depends on four inputs:

- The holdout's expected outcome rate over the fixed window.
- The smallest absolute or relative lift worth paying to detect.
- The chance of a false positive that the team will accept.
- The power required to detect the chosen lift if it is real.

Count accounts, not contacts. A list with 5,000 contacts across 900 companies has 900 assignment units for this test.

An even treatment and holdout split gives the most precision for a fixed audience. A smaller holdout preserves more sends but requires more total accounts for the same power. Use the [holdout size calculator](https://trysincerely.com/tools/holdout-size) before launch. Enter a credible baseline and the smallest result that would change the decision. Do not enter an optimistic lift just to make the required sample smaller.

If the eligible audience is too small, keep the holdout and accept a wider interval. Repeat the same design across comparable cohorts, then pool them under a written rule. The [small-audience measurement guide](https://trysincerely.com/guides/direct-mail-small-audiences) explains when that is defensible. Removing the holdout produces a bigger send, not better evidence.

## Preserve intention to treat

Analyze each account in the group assigned at launch, whether or not execution went perfectly. This is [intention to treat](https://trysincerely.com/glossary/intention-to-treat).

Suppose a treatment account has a bad address, a seller misses a call, or the recipient never opens the email. Moving that account out of treatment would remove hard cases after randomization and make the program look better than the program the team can actually run. Keep it in treatment. Operational failure is part of the effect of assigning the real workflow.

A per-protocol view can help diagnose delivery. Label it as a secondary operational analysis. It cannot replace the intention-to-treat result because completion was not randomly assigned.

## Watch for contamination

Contamination narrows the difference between groups and makes the test answer a muddier question. Common forms include:

- A rep sends a one-off touch to a holdout account.
- Different contacts at one account land in different groups.
- Paid media or an event targets only one arm without being part of the treatment.
- Accounts enter or leave the eligible list after assignment.
- A control account receives the physical piece because two systems use different identifiers.
- CRM stages are backfilled without the date the account actually crossed the threshold.

Do not quietly delete contaminated accounts. Log the event, keep the account in its assigned group, and report the contamination rate. If it is large, the result estimates the effect of the program as operated, leakage included. That may be less flattering and more useful.

## Instrument the operating path

The readout is only as trustworthy as the ledger behind it. Capture these records before asking analytics to reconstruct them.

| Record               | Minimum fields                                                                   | Why finance needs it                                                     |
| -------------------- | -------------------------------------------------------------------------------- | ------------------------------------------------------------------------ |
| Eligibility snapshot | Account ID, qualification reason, prior stage, segment, snapshot time            | Shows who could enter the test and blocks selection after results appear |
| Assignment ledger    | Account ID, campaign ID, group, method, seed or salt, assignment time            | Proves the comparison was randomized and reproducible                    |
| Treatment ledger     | Planned touch, completed touch, sender, channel, timestamp, delivery state       | Separates a weak idea from failed execution                              |
| Outcome ledger       | Account ID, primary stage event, event time, source system, prior value          | Applies the same outcome rule to both groups                             |
| Cost ledger          | Media, data, production, postage, tools, and seller time included by policy      | Makes the economic denominator explicit                                  |
| Exception ledger     | Suppression, address failure, rep override, leakage, duplicate, and cancellation | Exposes contamination and operational loss                               |

Use stable account IDs across the CRM, engagement systems, and fulfillment records. Store event timestamps rather than only the latest state. A current CRM stage cannot tell you whether the account entered pipeline inside the test window.

Audit a sample of both groups before launch and again before the readout. Check that the groups share the same eligibility rule, sales coverage, and outcome collection. Randomization cannot repair a different measurement rule for the holdout.

## A worked example

This example is illustrative, not a customer result or an outbound benchmark.

A team freezes 2,000 eligible accounts and assigns 1,000 to its outbound program and 1,000 to holdout. The primary outcome is a new sales-accepted opportunity within 90 days. When the window closes:

- Treatment has 89 opportunities, an 8.9% rate.
- Holdout has 70 opportunities, a 7.0% rate.
- Absolute lift is 1.9 percentage points.
- Relative lift is 27.1% against the holdout baseline.
- The point estimate is 19 incremental opportunities across the 1,000 treatment accounts.

The approximate 95% confidence interval for absolute lift runs from -0.47 to 4.27 percentage points. In account counts, the data support a range from about 5 fewer to 43 more opportunities. The two-sided p-value is 0.116. This test does not prove that outbound created pipeline, even though the midpoint is positive.

The [lift significance calculator](https://trysincerely.com/tools/lift-significance) produces those figures from the four account counts. The American Statistical Association's [statement on p-values](https://www.amstat.org/asa/files/pdfs/p-valuestatement.pdf) explains why a p-value does not measure effect size or business importance. Finance needs the lift, interval, assumptions, and cost, not a green or red significance badge by itself.

## Turn the result into a finance decision

Do not value every dollar of attributed pipeline as if it were revenue. Convert the incremental opportunity estimate into expected gross profit using assumptions finance already owns.

For the illustrative test, assume:

- Average first-year contract value is $50,000.
- The win rate from a sales-accepted opportunity is 20%.
- Gross margin is 80%.
- The fully loaded outbound test costs $40,000.

Expected gross profit per opportunity is `$50,000 x 20% x 80% = $8,000`.

At the 19-opportunity midpoint, expected incremental gross profit is `$152,000`. Subtract the $40,000 program cost and the midpoint net benefit is $112,000, or 280% of cost. That is a planning estimate, not a proven return. The confidence interval still includes harm, so the downside case is a loss.

Give finance four numbers together:

1. The attributed pipeline total, clearly labeled as attribution.
2. The incremental opportunity estimate and its confidence interval.
3. Expected gross profit under finance-approved win-rate and margin assumptions.
4. Fully loaded cost and the decision that follows.

A useful decision may be to repeat the same test until the interval rules out an uneconomic result. Another may be to stop because even the top of the interval cannot clear the required return. The test should make that boundary visible before the team spends more.

## How Sincerely fits

Sincerely can test the incremental value of physical mail inside an outbound program. Keep the digital sequence the same for both groups, assign whole accounts to mail or holdout, lock the primary outcome and window at launch, and compare CRM outcomes after the window closes. The product keeps failed addresses and missed fulfillment in the assigned arm, then reports lift with an interval.

Replies, scans, and postal events remain useful operating signals. They do not replace the holdout. A postal delivery event confirms a delivery point, not recipient attention.

## Sources and methodology

This guide applies established experiment-reporting principles to B2B outbound. It does not claim a public B2B outbound benchmark.

- [NIST's design of experiments handbook](https://www.itl.nist.gov/div898/handbook/pri/section1/pri11.htm) supports defining the objective, response, factors, and experimental design before execution.
- [CONSORT 2025](https://www.bmj.com/content/389/bmj-2024-081123.full) is a current reporting guideline for randomized trials. Its fields differ from B2B marketing, but its requirements to report allocation, outcomes, participant flow, analysis, and harms inform the audit trail described here.
- [The CONSORT cluster extension](https://www.bmj.com/content/345/bmj.e5661) explains contamination, group-level assignment, and the need to analyze at the randomized level. This guide maps the cluster to a company account because contacts at one company share an opportunity.
- [The American Statistical Association's p-value statement](https://www.amstat.org/asa/files/pdfs/p-valuestatement.pdf) supports reporting effect size and uncertainty rather than treating a threshold as proof or business importance.
- Sincerely's [holdout size calculator](https://trysincerely.com/tools/holdout-size) uses a two-proportion design with a configurable allocation and power. Its [lift calculator](https://trysincerely.com/tools/lift-significance) compares treatment and holdout rates, reports absolute and relative lift, and gives a 95% confidence interval.

The 2,000-account example and every financial assumption are hypothetical. The arithmetic shows how to report a test whose point estimate is positive but whose interval remains inconclusive.

## Related questions

### Can influenced pipeline prove outbound ROI?

No. Influenced pipeline applies a credit rule to accounts that converted. It does not estimate what would have happened without outbound. Use it as an operating view, then use a randomized holdout for the causal claim. [Attribution versus incrementality](https://trysincerely.com/guides/attribution-vs-incrementality) shows the difference with the same campaign counts.

### What should the holdout receive?

It should receive the comparison condition written into the test. To measure the whole outbound program, compare it with business as usual. To measure the added value of one channel, keep the other channels the same and withhold only that channel.

### What if sales will not hold out half the accounts?

Use a smaller holdout and calculate the larger sample it requires. Show the precision cost before choosing the split. A small holdout is still useful when it is large enough for the planned effect.

### Should pipeline dollars be the primary outcome?

Usually not for an early test. A few large deals can dominate the result. Start with a strict account-level opportunity event, then report pipeline value and eventual gross profit as secondary outcomes with their own uncertainty.

### Can you remove accounts that never received the treatment?

Not from the primary analysis. Keep them in the group assigned at launch. Report delivery failures and missed touches separately so the team can improve execution without biasing the causal estimate.

### How do you prove outbound for a small named-account list?

You may not be able to prove it in one cohort. Keep the account holdout, report the wide interval, and pool only later cohorts that preserve the same eligibility rule, treatment, outcome, and window. Do not substitute matchback for missing power.

## Related questions

- [Attribution vs incrementality, credit versus cause](https://trysincerely.com/guides/attribution-vs-incrementality): Attribution assigns credit for conversions that happened. Incrementality estimates which conversions would not have happened without the campaign. A worked example shows why the two numbers routinely disagree.
- [Direct mail holdout and sample size calculator](https://trysincerely.com/tools/holdout-size): Calculate the sample size a direct mail holdout needs from your baseline rate and minimum detectable lift, at 80 or 90% power, with the real cost of an unequal split.
- [Direct mail incrementality and lift calculator](https://trysincerely.com/tools/lift-significance): Compare mailed and holdout conversion rates to get absolute and relative lift, a 95% confidence interval, statistical significance, and how many conversions were truly incremental.
- [How to measure direct mail with a small audience](https://trysincerely.com/guides/direct-mail-small-audiences): Keep an account-level holdout, predeclare one outcome, and pool comparable cohorts when a single B2B direct mail campaign is too small for a powered result.

---

Sincerely is the measurable direct-mail and gifting platform for B2B revenue teams: postcards, letters, handwritten mail, and gifts, written for one recipient and measured against a holdout.

Contact Sincerely: https://trysincerely.com/contact

Agent routing index: https://trysincerely.com/llms.txt
