RFM — recency, frequency, monetary value — has lasted decades for good reasons. It is easy to explain, cheap to compute, needs nothing but order history, and it is dramatically better than treating your customer file as one undifferentiated list.
If you are not segmenting at all, RFM is the right next step. This is about what happens after that, because most brands stay on RFM long past the point where it is holding them back.
What RFM gets right
Credit where it is due. RFM captures three genuinely predictive dimensions, it produces segments a marketing team can actually reason about, and it is transparent — when a customer lands in "at risk," you know exactly why.
That transparency is worth a lot, and it is why RFM should stay in your reporting even after you outgrow it for targeting.
Where it breaks
It describes the past
Every RFM dimension is historical. "Last ordered 94 days ago, has ordered 6 times, has spent $840" is a complete description of what already happened and contains no forward statement.
By the time recency puts someone in an at-risk bucket, they have already stopped buying. You are not predicting churn; you are recording it. The whole value of prediction is acting during the window when behavior is still changeable, and RFM has no access to that window.
Everyone is on the same clock
This is the deepest flaw. RFM recency uses one scale for your entire base — so 90 days means the same thing for a customer who reorders every three weeks and one who reorders twice a year.
For the first, 90 days is a four-cycle emergency. For the second it is completely normal. RFM puts them in the same bucket and you send them the same email. One is long gone; the other is baffled to hear you miss them.
Any serious churn model normalises to each customer's own cadence, which is the single biggest improvement available over RFM.
It cannot see direction
RFM sees levels, not trajectories. Consider two customers, both with six orders and both 40 days since their last:
- Customer A's gaps have run 30, 32, 29, 31, 33 days. They are 40 days out — slightly late, nothing alarming.
- Customer B's gaps have run 18, 24, 31, 38, 44 days. Their interval has more than doubled. They have been leaving for months.
Identical RFM scores. Completely different situations. The information that distinguishes them — the trend in the gaps — is exactly what RFM discards.
It has no idea why
RFM knows nothing about a nine-day delivery, a two-star review, an unresolved support ticket or a failed subscription charge. Those are the actual causes of churn, and they are invisible to a model built from three order-history aggregates.
Which means RFM can sometimes tell you that someone is disengaging but never why — and without the reason you cannot choose a message. You default to a discount, which is the wrong response to most of these causes.
The boundaries are arbitrary
RFM buckets are usually quintiles or hand-picked thresholds. A customer at the 79th percentile and one at the 81st get different treatment despite being nearly identical, and nobody can defend where the line sits. A probability output does not have this problem — it is continuous, and you can set action thresholds deliberately based on what you intend to do.
What a behavioral model adds
The move is not from three variables to three hundred for its own sake. It is from levels to patterns, and from description to prediction.
- Cadence-relative recency. Overdue measured against each customer's own interval rather than a global number.
- Trend features. Gap acceleration, order-value slope, changes in discount reliance — the direction of travel, not just the position.
- Lifecycle features. First-to-second-order gap, what happened in the first 30 and 60 days, which product they entered on.
- Experience features. Shipping performance, review sentiment, support outcomes, payment failures.
- Learned weights. Rather than you deciding that recency matters more than frequency, the model learns the weighting from what actually preceded churn in your data.
That last point is underrated. RFM implicitly assumes the three dimensions matter equally, or in whatever ratio you choose. For some catalogues frequency is far more predictive than recency; for others the reverse. A model fitted to your history does not need you to guess.
The fair objections
"Models are black boxes." They are if you build them that way. A retention system should surface the drivers behind every score in plain language — slowing cadence, shipping delays, subscription paused. If it cannot explain itself, that is a product failure, not an inherent property of modelling.
"We do not have enough data." Sometimes true, and the honest answer is that below roughly a hundred repeat customers and a year of history, a good heuristic beats a trained model. A well-built system detects this and falls back rather than overfitting. But note that the heuristic worth falling back to is still cadence-relative — it is not RFM.
"RFM works fine for us." Compared to nothing, certainly. The question is what it costs you: the customers whose cadence made them invisible to a global recency rule, and the margin spent discounting people who needed something other than a discount. Both are measurable with a holdout test.
Keep RFM for what it is good at
This is not an argument to delete your RFM segments. It remains excellent for reporting, for describing your base to people who do not want a probability distribution, and for sanity-checking a model's output — if your model says your best RFM segment is high risk, something is wrong somewhere.
Use RFM to understand your customer base. Use a behavioral model to decide who to contact today, and why.
Telltale scores every customer against their own history with the drivers attached, and falls back to transparent heuristics when there is not enough data to do better. Install from the Shopify App Store.
