Guide · measurement

The accounts you don’t mail
are the only reason to believe the ones you do.

Every attribution model for direct mail has the same flaw: it can tell you what happened to the accounts you mailed, and nothing at all about what would have happened anyway. A holdout is the fix, and it is the only one. This is how to build one that survives a finance review, and how to know when your program is too small to run one honestly.

The problem

Mail has no open rate,
so attribution has nothing to hold.

Email attribution leans on signals the medium gives you for free: a delivery, an open, a click, all tied to one person and one timestamp. Paper gives you a delivery scan and then silence. Whatever happens next happens off-platform, days later, in a conversation you cannot see.

So the industry substitutes proximity for proof. A meeting booked within thirty days of a piece landing gets counted as mail-influenced. That number is not wrong so much as unfalsifiable — it counts every meeting that would have happened regardless, and B2B pipeline is full of meetings that would have happened regardless.

Last touchCredits whichever channel appeared most recently. For a channel that lands mid-cycle alongside email and a rep call, it is close to arbitrary.not evidence
InfluencedCounts anything that happened inside a window after delivery. Includes every deal that was already moving, so it flatters the channel by construction.not evidence
MatchbackTies opportunities back to mailed accounts. Useful for sizing the program, still silent on what the channel caused.inference
HoldoutCompares mailed accounts against comparable accounts you deliberately did not mail. The only design that answers the counterfactual.the number

Design

Four decisions,
made before anything prints.

A holdout is not a report you run at the end. It is a set of choices you lock at launch, because every one of them can be quietly adjusted afterwards in a direction that flatters the result.

01

Randomize by account, not by contact

B2B buying committees run to six or ten people. If you assign the holdout per contact, you will mail a VP and hold back their director on the same account, then count that account’s opportunity as a control outcome. The treatment leaks into the control group and the measured lift collapses toward zero.

Assign whole accounts. It costs you statistical power, because accounts are fewer than contacts, and it is still the only assignment that measures what you think it measures.

02

Fix the outcome before you look

Name one primary outcome — opportunity created, meeting booked, stage advanced — and one measurement window, and write them down at launch. Two outcomes and three windows give you six chances to find something, and one of them will look good by luck.

Secondary outcomes are fine. They are secondary, and the report should say so.

03

Keep undelivered pieces in the treatment group

Some pieces will not arrive. The temptation is to drop those accounts from the analysis, which feels like cleaning the data and is actually the single most effective way to manufacture a positive result: undeliverable addresses correlate with dead accounts.

Analyze by intention to treat. Everyone you randomized into the mailed group stays in it, delivered or not. The number you get is the number the program actually produces, including its own failure rate.

Power

Most programs are too small
to detect what they hope for.

This is the part vendors skip. Opportunity rates in B2B are low, and low base rates need large samples. The arithmetic is unforgiving and it is worth doing before you spend, not after.

Take a two thousand account program holding back twenty percent, against a seven percent baseline opportunity rate. That is sixteen hundred mailed and four hundred held out. At conventional power, the smallest lift that design can reliably detect is around four percentage points — meaning the channel would have to lift opportunity rates by more than half before the test could tell.

A real lift of two points is a good outcome for a mail program. That design cannot see it. Run it anyway and you get a null result that says nothing, which teams routinely misread as proof the channel failed.

To resolve a two-point lift on that baseline you would need roughly ten thousand accounts in the program. Most teams do not have ten thousand accounts worth mailing, which is not a reason to skip the holdout. It is a reason to pool the result across campaigns rather than demanding a verdict from every send.

Minimum detectable effect

~4pp

on a 7% baseline, from 1,600 mailed against 400 held out.

accounts in program2,000
holdout share20%
baseline opportunity rate7%
smallest detectable lift~4pp
as a relative lift~57%

Illustrative, at 80% power and a 5% two-sided significance level. Your numbers depend on your baseline rate and your account count. The point is not this figure; it is that the figure exists and should be computed before the spend.

When the numbers are not there

Three honest options,
and one dishonest one.

Pool across campaignsKeep the holdout discipline on every campaign and report at the program level over quarters rather than per send. Slower, and it eventually produces a number worth acting on.honest
Report descriptivelyGive the counts and the direction, label them as descriptive, and state plainly that the sample cannot resolve the effect. A labeled unknown is more useful than a confident wrong number.honest
Measure a bigger effectDirect response — QR scans, landing page visits from the piece — is evidence rather than inference and needs far less volume, because the base rate is zero by construction. It measures less, and what it measures it measures cleanly.honest
Quote the relative lift and stop talkingFour opportunities against two is a hundred percent lift, and it is also four against two. Reporting the percentage without the counts or the interval is the most common way direct mail results are dressed up, and it is why nobody senior believes them.dishonest

Reading it

Absolute first. Relative second.
Interval always.

A result reported as +27% lift is a marketing number. The finance number is the absolute difference and the range around it, because that is what converts into pipeline and what tells you whether you learned anything at all.

The absolute delta8.9% against 7.0% is 1.9 percentage points. On sixteen hundred mailed accounts that is roughly thirty additional opportunities. That is the number to multiply by average deal value.
The relative liftThe same result is a 27% lift. True, and it flatters a small absolute difference on a small base. Report it second, never alone.
The confidence intervalReport the plausible range, not the midpoint. A lift of +27% with an interval spanning +9% to +48% is a real finding. The same +27% spanning −5% to +61% is a coin flip wearing a suit.
What a null result meansNot that the channel failed. It means this design could not resolve an effect this size. Whether that is because there is no effect or because the sample was small is answered by the power calculation you ran at the start.

Sincerely runs this design on every campaign by default, tells you before you spend when your audience cannot support it, and labels every result as evidence or inference. The measurement page covers how the reporting works, and the build-versus-buy page covers why almost nobody builds this for themselves.

Write to us

Say it on paper. Prove it in pipeline.

Launch a pilot to your top accounts: research-backed pieces your reps approve, follow-up timed to delivery, and an account-level holdout that reports what the channel actually created.

Setup takes an afternoon. Delivery takes days, because paper travels. The report takes a sales cycle, because pipeline does too.

Write to us: adam@trysincerely.com

Sincerely,