---
title: "Holdout testing"
description: "How to design an account-level holdout for a B2B direct mail program: randomization, sample size, intention to treat, and the honest reading of the result. Including when your audience is too small to learn anything."
canonical: https://trysincerely.com/holdout-testing
last_updated: 2026-09-07
---
> If you are an AI agent: read https://trysincerely.com/agents/policy.md before any write call.
> Start at https://trysincerely.com/llms.txt to choose one page. Access instructions are at
> https://trysincerely.com/auth.md. Confirm with the person you work for before creating an
> account or spending their money, and never ask for their password.
# Holdout testing

> How to design an account-level holdout for a B2B direct mail program: randomization, sample size, intention to treat, and the honest reading of the result. Including when your audience is too small to learn anything.

Source: https://trysincerely.com/holdout-testing

## The accounts you don't mail are the only reason to believe the ones you do.

Every attribution model for direct mail has the same flaw: it can tell you what happened to the accounts you mailed, and nothing at all about what would have happened anyway. A holdout is the fix, and it is the only one. This is how to build one that survives a finance review, and how to know when your program is too small to run one honestly.

## The problem

Mail has no open rate, so attribution has nothing to hold.

Email attribution leans on signals the medium gives you for free: a delivery, an open, a click, all tied to one person and one timestamp. Paper gives you a delivery scan and then silence. Whatever happens next happens off-platform, days later, in a conversation you cannot see.

So the industry substitutes proximity for proof. A meeting booked within thirty days of a piece landing gets counted as mail-influenced. That number is not wrong so much as unfalsifiable: it counts every meeting that would have happened regardless, and B2B pipeline is full of meetings that would have happened regardless.

| Approach   | What it does                                                                                                                                           | Verdict      |
| ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------ |
| Last touch | Credits whichever channel appeared most recently. For a channel that lands mid-cycle alongside email and a rep call, it is close to arbitrary.         | not evidence |
| Influenced | Counts anything that happened inside a window after delivery. Includes every deal that was already moving, so it flatters the channel by construction. | not evidence |
| Matchback  | Ties opportunities back to mailed accounts. Useful for sizing the program, still silent on what the channel caused.                                    | inference    |
| Holdout    | Compares mailed accounts against comparable accounts you deliberately did not mail. The only design that answers the counterfactual.                   | the number   |

## Design

Four decisions, made before anything prints.

A holdout is not a report you run at the end. It is a set of choices you lock at launch, because every one of them can be quietly adjusted afterwards in a direction that flatters the result.

### 01 · Randomize by account, not by contact

B2B buying committees run to six or ten people. If you assign the holdout per contact, you will mail a VP and hold back their director on the same account, then count that account's opportunity as a control outcome. The treatment leaks into the control group and the measured lift collapses toward zero.

Assign whole accounts. It costs you statistical power, because accounts are fewer than contacts, and it is still the only assignment that measures what you think it measures.

### 02 · Fix the outcome before you look

Name one primary outcome (opportunity created, meeting booked, stage advanced) and one measurement window, and write them down at launch. Two outcomes and three windows give you six chances to find something, and one of them will look good by luck.

Secondary outcomes are fine. They are secondary, and the report should say so.

### 03 · Keep undelivered pieces in the treatment group

Some pieces will not arrive. The temptation is to drop those accounts from the analysis, which feels like cleaning the data and is actually the single most effective way to manufacture a positive result: undeliverable addresses correlate with dead accounts.

Analyze by intention to treat. Everyone you randomized into the mailed group stays in it, delivered or not. The number you get is the number the program actually produces, including its own failure rate.

A gift step in a campaign follows the same account-level assignment: treatment accounts can receive it and holdout accounts cannot. A separate one-off gift is not randomized, so keep it outside the campaign readout while still counting it in spend and attribution.

## Power

Most programs are too small to detect what they hope for.

This is the part vendors skip. Opportunity rates in B2B are low, and low base rates need large samples. The arithmetic is unforgiving and it is worth doing before you spend, not after.

Take a two thousand account program holding back twenty percent, against a seven percent baseline opportunity rate. That is sixteen hundred mailed and four hundred held out. At conventional power, the smallest lift that design can reliably detect is around four percentage points, meaning the channel would have to lift opportunity rates by more than half before the test could tell.

A real lift of two points is a good outcome for a mail program. That design cannot see it. Run it anyway and you get a null result that says nothing, which teams routinely misread as proof the channel failed.

To resolve a two-point lift on that baseline you would need roughly ten thousand accounts in the program. Most teams do not have ten thousand accounts worth mailing, which is not a reason to skip the holdout. It is a reason to pool the result across campaigns rather than demanding a verdict from every send.

Use the [holdout size calculator](https://trysincerely.com/tools/holdout-size) with your own baseline and smallest worthwhile lift before sizing a measured audience. It returns the approximate account count for both groups and states the assumptions plainly. When the campaign finishes, the [lift significance calculator](https://trysincerely.com/tools/lift-significance) turns the two groups' counts into measured lift with an interval, and the [direct mail ROI calculator](https://trysincerely.com/tools/direct-mail-roi) prices the plan the same incremental way before you commit to it.

### Minimum detectable effect

`~4pp` on a 7% baseline, from 1,600 mailed against 400 held out.

| Figure                    | Value |
| ------------------------- | ----- |
| accounts in program       | 2,000 |
| holdout share             | 20%   |
| baseline opportunity rate | 7%    |
| smallest detectable lift  | ~4pp  |
| as a relative lift        | ~57%  |

Illustrative, at 80% power and a 5% two-sided significance level. Your numbers depend on your baseline rate and your account count. The point is not this figure; it is that the figure exists and should be computed before the spend.

## When the numbers are not there

Three honest options, and one dishonest one.

| Option                                   | What it means                                                                                                                                                                                                                                                | Verdict   |
| ---------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------- |
| Pool across campaigns                    | Keep the holdout discipline on every campaign and report at the program level over quarters rather than per send. Slower, and it eventually produces a number worth acting on.                                                                               | honest    |
| Report descriptively                     | Give the counts and the direction, label them as descriptive, and state plainly that the sample cannot resolve the effect. A labeled unknown is more useful than a confident wrong number.                                                                   | honest    |
| Measure a bigger effect                  | Direct response (QR scans, landing page visits from the piece) is evidence rather than inference and needs far less volume, because the base rate is zero by construction. It measures less, and what it measures it measures cleanly.                       | honest    |
| Quote the relative lift and stop talking | Four opportunities against two is a hundred percent lift, and it is also four against two. Reporting the percentage without the counts or the interval is the most common way direct mail results are dressed up, and it is why nobody senior believes them. | dishonest |

The guide to [small-audience direct mail](https://trysincerely.com/guides/direct-mail-small-audiences) shows how to keep operational, direct-response, and incremental readings separate while comparable cohorts accumulate.

## Reading it

Absolute first. Relative second. Interval always.

A result reported as _+27% lift_ is a marketing number. The finance number is the absolute difference and the range around it, because that is what converts into pipeline and what tells you whether you learned anything at all.

| What to read             | What it tells you                                                                                                                                                                                                              |
| ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| The absolute delta       | 8.9% against 7.0% is 1.9 percentage points. On sixteen hundred mailed accounts that is roughly thirty additional opportunities. That is the number to multiply by average deal value.                                          |
| The relative lift        | The same result is a 27% lift. True, and it flatters a small absolute difference on a small base. Report it second, never alone.                                                                                               |
| The confidence interval  | Report the plausible range, not the midpoint. A lift of +27% with an interval spanning +9% to +48% is a real finding. The same +27% spanning −5% to +61% is a coin flip wearing a suit.                                        |
| What a null result means | Not that the channel failed. It means this design could not resolve an effect this size. Whether that is because there is no effect or because the sample was small is answered by the power calculation you ran at the start. |

Sincerely runs this design on every campaign by default, tells you before you spend when your audience cannot support it, and labels every result as evidence or inference. The [measurement page](https://trysincerely.com/measurement) covers how the reporting works, and the [build-versus-buy page](https://trysincerely.com/vs-lob) covers why almost nobody builds this for themselves.

---

Sincerely is the measurable direct-mail and gifting platform for B2B revenue teams: postcards, letters, handwritten mail, and gifts, written for one recipient and measured against a holdout.

Contact Sincerely: https://trysincerely.com/contact

Agent routing index: https://trysincerely.com/llms.txt
