How do you measure recommendation quality beyond click-through rate?
Click-through rate tells you a recommendation was tempting, not that it was good. The metrics that predict revenue are conversion per recommendation view, revenue per session, catalog coverage, and lift proven against a holdout group.
Why click-through rate lies
Click-through is easy to game and easy to misread. A row of deeply discounted clearance items gets enormous click-through and teaches you nothing about recommendation quality. A "customers also bought" row full of cheap add-ons gets clicks that never become orders. Worse, optimizing for clicks pushes the engine toward clickbait: the most clickable products are often the cheapest, the most familiar, or the most discounted, which is the opposite of what a good recommendation should surface. CTR is a diagnostic, like a temperature reading. Useful, but nobody runs a store on temperature alone.
The revenue metrics that matter
Start with conversion per recommendation view: of the shoppers who saw the row, how many bought something from it. This punishes clickbait automatically, because a click that does not convert counts against the row. Next, revenue per session for sessions that interacted with recommendations versus sessions that did not, controlled for the fact that engaged shoppers were already more likely to buy. The clean version is a holdout test: a small percentage of traffic sees a neutral default, like bestsellers, and you compare revenue per session between the groups. That single number, holdout-tested lift, is the one to put in front of leadership. Everything else is supporting detail.
Catalog coverage and diversity
A recommendation engine that only ever surfaces 5 percent of your catalog is a bestseller list with extra steps. Track coverage: what share of your catalog appears in recommendations over a month. Low coverage means new products, long-tail items, and high-margin niche products never get a chance, which is exactly the inventory a recommendation engine should be unlocking. Track diversity per session too. If every row on the page shows variations of the same product, the engine is not personalizing, it is repeating. A healthy engine spreads attention across categories while staying relevant to the shopper.
Position and cannibalization
Two rows that look identical in the dashboard can behave very differently depending on where they sit. Measure each recommendation placement separately, because a row above the fold and a row in the footer are different products from a measurement perspective. And watch for cannibalization: a new row that converts well might just be stealing clicks from the row next to it. The test is total page revenue with and without the new row, not the new row's metrics in isolation. This is where holdout testing earns its keep a second time: test placements, not just engines.
A practical measurement stack
For a store getting started, this is enough:
- Holdout lift in revenue per session: the headline number, reviewed monthly.
- Conversion per recommendation view: per placement, reviewed weekly, to catch clickbait drift.
- Catalog coverage: monthly, to make sure the long tail is being served.
- Click-through rate: as a diagnostic when one of the above moves unexpectedly.
Notice the order. CTR is last, and only as a debugging tool. The stores that measure recommendations well all end up here: revenue first, behavior second, clicks last.
Reviewed
Published Sep 24, 2026.