The Loyalty Illusion
Loyalty programs are among the largest recurring customer-spend lines in retail, travel, financial services, hospitality and gaming — and their performance is routinely reported as the spending of the people who joined them. Member spend is a description of who enrolled. It is not proof that the program caused anything.
“Loyalty is not what customers spend. It is what your action changes.”
This report sets out the causal question a loyalty program has to answer, the four distinct economic ledgers hiding inside most programs, and the decision system required to tell rewarded behavior apart from changed behavior.
The Loyalty Illusion
Are loyalty programs creating incremental customer value — or just rewarding behavior that would have happened anyway?
Figures are labelled by evidence class. Published study results describe the category, market and period studied; they are cited as evidence of a mechanism, not as transferable constants. No universal loyalty incrementality coefficient is claimed anywhere in this research.
Five conclusions that change how a program should be judged.
Programs recruit the customers who already buy the most. Comparing members to non-members therefore measures selection first and program effect second — if at all. Reported member share of revenue describes who enrolled, not what enrollment caused.
Within any program some customers are moved by a reward, some would have purchased regardless, and some react negatively to intrusive contact. A single average uplift number hides those populations and reliably funds the wrong one.
A behavior-change program, a currency and partner ecosystem, and a paid-access subscription earn money in structurally different ways. Judging all three against the same lift metric misprices at least two of them.
A churn model ranks who is likely to leave. A decision system estimates who is saveable by a specific action. The two rankings routinely disagree, and only the second one can justify spending a reward.
A one-off test dates quickly as customers, competitors and offers change. Incrementality has to be re-estimated on a standing cadence, with holdouts preserved deliberately rather than sacrificed to short-term reach.
Revenue lift is not value. The counterfactual is.
A point issued is a cost recognised now against behavior that may or may not have required it. The relevant quantity is not what a rewarded customer spent, but the difference between what they spent and what they would have spent untreated — Y(1) − Y(0). Observational reporting supplies Y(1) in abundance and Y(0) never.
Because programs recruit heavy buyers first, the untreated comparison group is systematically different from the treated one. That is selection, not effect. The Leenheer grocery study is instructive precisely because correcting for self-selection reduced the apparent program effect by roughly seven times — in that context, with that data.
A maintained holdout is the price of knowing. Without a deliberately untreated control, a loyalty P&L reports activity and calls it return.
Three business models, wearing one name.
Behavior-change programs, currency and ecosystem programs, and paid-access memberships earn money in structurally different ways. The same lift metric cannot price all three.
Rewards intended to alter purchase frequency, basket size or retention. Value exists only where the treated outcome differs from the untreated counterfactual, which means holdouts are not optional.
Points sold to partners create a real revenue and liability structure largely independent of any behavioral lift. Here the economics are breakage, issuance margin, redemption cost and partner mix.
Subscription and membership tiers are priced access, not persuasion. The relevant questions are willingness to pay, service cost to serve, and whether the tier changes behavior beyond the fee itself.
Two questions that look identical and are not.
Ranks customers by risk. Its top of list is often dominated by people who are leaving for reasons no offer addresses — where reward spend produces cost and no change.
Ranks customers by estimated treatment effect. It is the only ranking that maps onto a spending decision, because it is the only one that references the action.
Baseline value against signed uplift.
Two axes, four states, four different actions. Value alone tells you who matters; signed uplift tells you who your action can move. Only the intersection justifies spend.
Valuable customers who respond to the action. This is the only cell where increased reward spend is straightforwardly defensible.
Valuable, but not moved by the treatment. Recognition and service quality — not incremental reward cost, which buys behavior already occurring.
Responsive but currently small. Worth measured investment where the estimated incremental contribution clears the reward cost.
Treatment produces nothing, or actively harms the relationship. Suppression protects both margin and the customer experience.
Some customers respond worse when treated — over-contacted, discount-trained or reminded of a decision they had not been making. Averaging uplift across a base hides that population entirely, which is why the effect must be carried signed, per customer, and never as a single programme-level mean.
Six outcomes, one test.
Every mechanic in the program resolves to one of these, judged on incremental customer lifetime value net of reward and servicing cost — never on observed value among the treated.
Incremental CLV clearly exceeds reward and servicing cost, with a stable estimate across periods.
Aggregate effect is positive but concentrated — target the responsive population rather than the whole base.
The mechanic is understood and the effect is weak. Change the offer structure before changing the budget.
The estimate is too noisy or too new to authorise. Run a controlled design before committing spend.
Positive but marginal returns. Lower reward intensity and re-measure rather than defending the current level.
No credible incremental contribution after adequate measurement. The spend is buying existing behavior.
Eight steps, always in this order.
The order is the method. Most loyalty programs fail measurement by scaling before they instrument, or by targeting propensity before they have estimated response.
Five numbers that survive scrutiny.
Contribution measured against a maintained holdout, not against non-members or prior-year members.
Per-customer treatment effect with its sign preserved, so negative responders are visible instead of averaged away.
Lifetime value attributable to the action, not lifetime value observed among the treated.
The efficiency ratio that makes reward budgets comparable across mechanics and segments.
Issuance margin, redemption behavior, breakage and the balance-sheet position the currency creates.
The report includes a fully illustrative policy simulation across a hypothetical 5-million-member base, comparing blanket treatment with uplift-optimized targeting. Under its stated assumptions the two policies differ by roughly $29.9M. This figure is a worked demonstration of how targeting policy changes economics — it is not an industry estimate, a benchmark, or a client outcome.
Nine chapters across thirty-seven pages.
What a reward costs, what it defers, and why revenue lift is not the same as value created.
Y(1) − Y(0), and why member spend answers only half of it.
Behavior change, currency, partner economics and paid access, priced separately.
Who is likely to churn versus who is saveable by this action.
Baseline value against signed uplift, and the four actions it implies.
Holdouts, treatment assignment and the designs that survive an audit.
Rebuilding lifetime value so it reflects the action rather than the population.
Audit through conditional scale, in the order that keeps each step honest.
What the evidence base can and cannot support, stated before the conclusions.
Built for the people who fund the program.
A defensible basis for what the program is actually earning, mechanic by mechanic.
Targeting on estimated response to an action rather than on propensity to purchase anyway.
A P&L split that separates behavioral lift from currency and partner economics.
The move from predictive scoring to causal estimation, with the designs that support it.
Whether loyalty spend is compounding customer value or subsidising existing behavior.
What this research does not claim.
The limitations are stated before the conclusions, because the central argument of the report is about the difference between observation and evidence. It would be self-defeating to make that case and then overreach.
Disclosed loyalty revenue, deferred revenue and breakage describe financial architecture. They contain no randomized customer-level comparison, so they cannot establish incrementality.
Stated preference and satisfaction data describe how programs are experienced. They are not evidence of what a reward changed in behavior.
Published loyalty studies are anchored to a category, market and period. Their effect sizes travel poorly and are cited here as evidence of a mechanism, not as transferable constants.
This report proposes no global incrementality rate. Company-specific estimates require company-specific treatment and control data.
Get the full The Loyalty Illusion.
The complete publication — the economics of a point, the four ledgers, prediction versus causal decisioning, the uplift matrix, incremental CLV gates, the operating sequence, the illustrative policy simulation and the full methodology and limitations.
The research frames the question. Your own data estimates the effect.
No published study can tell you what your rewards changed. Incrementality is company-specific by construction: it requires your own customer-treatment and holdout data, one governed definition of contribution and lifetime value, and a measurement cadence that keeps the estimate current. Nucleus is where that layer sits.
External intelligence frames the decision problem. Internal intelligence supplies the counterfactual.
Explore Nucleus
Shemal authors the Dr.D Intelligence series and builds the analytics environments behind it. He works on governed metric layers, executive reporting and applied AI — which is why the research is written the way an operator would need it: definitions first, evidence labelled, conclusions stated plainly.
shemal@dr-danalytics.com