Here is a number almost every retention tool will show you: revenue attributed to your campaigns. Someone received an email, bought within a window, and the order gets counted.
It is a real number. It is also not the number you want, and the gap between the two is where a great deal of ecommerce marketing budget quietly disappears.
The problem with attributed revenue
Attribution answers: "how much revenue followed our campaign?" The question you actually care about is: "how much revenue happened because of our campaign?"
Those diverge badly in retention, and for a specific structural reason: retention marketing deliberately targets people who are likely to buy. That is the entire point of predictive targeting. You aim at customers inside their reorder window, or with high intent, or with an established habit.
Which means a large share of the people you contact were going to purchase whether or not you sent anything. When they do, attribution hands your campaign the credit.
The better your targeting gets, the worse this bias becomes. A perfectly-targeted campaign aimed at people about to reorder anyway will show spectacular attributed revenue and near-zero real impact. Optimising against attribution actively pushes you toward that failure mode — it rewards you for finding customers who need no persuasion.
What a holdout actually is
The fix is old, boring and borrowed from clinical trials: before you send, randomly hold back a slice of the customers who qualified for the campaign, and never contact them.
Everything else stays identical. Both groups had the same behavior, the same risk scores, the same eligibility. The only difference is that one got the message and one did not. Compare purchase rates afterwards, and the difference between them is your incremental lift — the part you caused.
A worked example. You have 10,000 customers qualifying for a reorder campaign. You hold back 10% at random:
- Contacted: 9,000 customers, 2,700 purchase — a 30% rate.
- Held back: 1,000 customers, 180 purchase — an 18% rate.
Attribution would report 2,700 recovered orders. The truth is that 18% of the contacted group — about 1,620 orders — would have happened anyway. Your real contribution is the 12-point difference: roughly 1,080 incremental orders.
That is a 40% haircut on the headline number. Uncomfortable, and far more useful. It is the number you can make decisions with.
Randomisation is the whole trick
The holdout must be chosen at random from within the qualifying audience. Every shortcut people take here destroys the comparison:
- Do not use "people we could not reach" as the control. Customers with no valid email are systematically different — usually older, less engaged, worse purchasers. You will overstate lift.
- Do not use a previous time period as the control. Seasonality, promotions and traffic mix all changed.
- Do not use non-qualifying customers. They are different by construction — that is why they did not qualify.
- Do not let the holdout receive the campaign through another route. If a held-back customer gets the same offer via a broadcast, the control is contaminated.
Sizing the holdout
There is a real cost here, and it should be stated plainly: customers in the holdout do not get contacted, so you forgo whatever incremental revenue you would have earned from them. That is the price of knowing whether the program works.
The trade-off is between statistical confidence and that forgone revenue. A very small holdout is cheap but produces noisy results you cannot act on; a large one gives tight estimates and costs more. Something in the range of five to ten percent of the eligible audience is a common, sensible starting point for most brands — large enough to detect a meaningful effect on campaigns of reasonable size, small enough that the cost is marginal.
What matters more than the exact percentage is that the holdout is persistent and consistently applied, so results accumulate over time rather than resetting every campaign.
Say "measuring" until you can say a number
Small campaigns produce noisy comparisons. If 40 people were contacted and 12 held back, the difference between the groups tells you almost nothing — a couple of orders either way swings the result completely.
The honest response is to withhold the number until it clears a reasonable confidence bar, and label it as still measuring. Reporting a precise-looking lift figure from a sample that cannot support it is worse than reporting nothing, because it will be believed and acted on.
Telltale reports results as "measuring" until they clear statistical verification, and — importantly — charges its performance fee only on lift that has cleared. If we cannot prove it, we do not bill for it.
What changes when you measure this way
Brands that switch from attribution to incrementality tend to discover three things, usually in this order:
Their best-performing campaign is not what they thought. The campaigns with the highest attributed revenue are often the ones aimed at the most-certain-to-buy customers — which is exactly where incremental lift is lowest. Meanwhile a campaign with modest attributed numbers turns out to be doing most of the real work.
Some campaigns have no effect at all. Every retention program has at least one send that is pure attribution theatre. Finding it saves money and list fatigue immediately.
Discounts are often unnecessary. Run the same campaign with and without an offer against a holdout and you frequently find the offer added cost without adding conversions. That single test can pay for a year of tooling.
Getting started
You do not need special software to begin. Pick your highest-volume retention campaign. Before the next send, randomly exclude 10% of the qualifying list. Wait for your normal purchase window. Compare purchase rates. Repeat for a few cycles before drawing conclusions.
The main thing standing between most brands and this practice is not difficulty — it is that holdouts make the headline numbers smaller, and nobody enjoys volunteering for that. It is still the right call. A smaller true number you can trust is worth more than a large one you cannot.
Telltale holds a control group back from every send automatically and reports lift against it, so incrementality is the default rather than a project. Install from the Shopify App Store, or read how this connects to predictive targeting.
