Skip to content
AI.info

Evaluation

Recommender Evaluation Beyond Relevance

Evaluate recommendations through relevance, coverage, diversity, novelty, calibration, long-term outcomes, and feedback-loop effects.

By the end you can

Reverse the ranking and the same users click differently

Logged interactions reveal responses to exposed items, not preferences over everything that could have been shown. A recommender that narrows exposure can reinforce its own evidence.

An eye-tracking study in 2005 measured how much of a click belongs to the item and how much to the slot it was put in. Participants searched Google through a proxy that silently rewrote the ranking before it reached the screen. There were 34 of them in the first phase, 22 in the second, and ten fixed questions. Swap the top two results and subjects still preferred link one, even when the second abstract had been judged more relevant. Reverse the whole list and the average rank of a clicked document moved from 2.66 to 4.03, while clicks per query fell from 0.80 to 0.64. Same documents, same queries, different positions, different clicks.

Joachims and colleagues ran it, and their own summary is the sentence to carry forward: “we conclude that clicks are informative but biased. While this makes the interpretation of clicks as absolute relevance judgments difficult, we show that relative preferences derived from clicks are reasonably accurate on average.”

It is not an artifact of one laboratory. Craswell and colleagues perturbed the ranking of a major search engine in 2008 and reported that “the probability of click is influenced by a document's position in the results page”. Different organisation, different method, same conclusion.

Evaluation therefore has to separate observed engagement from a question the log cannot answer: what would users have chosen under another ranking?

Reverse the list and the clicks move: mean clicked rank 2.66 to 4.03, clicks per query 0.80 to 0.64.

Comparison

Terms that are often collapsed

Each property has to be defined on its own, and these definitions have authors. Intra-list similarity — the measure behind every diversity number a recommender team reports — was introduced in 2005 by Ziegler and colleagues.

The evidence came in two parts. Offline: 361,349 BookCrossing ratings, condensed from a four-week crawl of 278,858 members and 1,157,112 ratings over 271,379 distinct ISBNs. Online: a three-week survey in which 2,125 users participated. Topic diversification lowered accuracy in every configuration. What it did to the list as a whole depended on the algorithm. It “has no measurable effects on user-based CF, but significantly improves item-based CF performance for diversification factors around 40%”.

The abstract states the trade-off without softening it: “Though being detrimental to average accuracy, we show that our method improves user satisfaction with recommendation lists, in particular for lists generated using the common item-based collaborative filtering algorithm.” The broader conclusion is the reason this section exists at all: “the user's overall liking of recommendation lists goes beyond accuracy and involves other factors, e.g., the users' perceived list diversity.”

So the sub-item below reading “may harm relevance” is not a caution. It is a measured cost that 2,125 people were nonetheless willing to pay — for one family of algorithms and not another.

And “depends on similarity measure” is literal, down to the sign of the quantity. A later survey by Castells and Jannach restates the average-pairwise-dissimilarity formula, then notes in a footnote that “Ziegler et al. [2005] and other authors define intra-list similarity instead of dissimilarity”. Two papers can report diversity and mean opposite directions.

FigureComparison · 4 columns

Intra-list diversity

Items within one list differ according to a chosen representation.

  • Can reduce redundancy
  • Depends on similarity measure
  • May harm relevance
  • Does not imply novelty

Novelty

Items are less familiar or globally popular to the user.

  • Can broaden discovery
  • Depends on exposure history
  • May surprise negatively
  • Not automatically useful

Serendipity

Relevant discovery is both useful and unexpected relative to a baseline.

  • Needs an expectation model
  • Hard to label offline
  • User-specific
  • Not identical to randomness

Catalog coverage

A broader fraction of items receive recommendation exposure.

  • Matters for suppliers
  • Can reveal concentration
  • Needs quality guardrails
  • Can be manipulated by low-value exposure

Example

A podcast experiment on Spotify with conflicting wins

Spotify ran a randomized field experiment on its own recommender in 2020. Some users were given personalized podcast recommendations. The rest were given the most popular podcasts among users in the same demographic group. Engagement and diversity were both measured off that one random assignment. That is why the two results cannot be argued away as different populations or different weeks.

The finding, in the authors' words: “We find that, on average, the treatment increased podcast streams by 28.90%. However, the treatment also decreased the average individual-level diversity of podcast streams by 11.51%, and increased the aggregate diversity of podcast streams by 5.96%, indicating that personalized recommendations have the potential to create patterns of consumption that are homogenous within and diverse across users, a pattern reflecting Balkanization.”

That is the abstract of the authors' working paper, by Holtz, Aral and colleagues. Spotify Research's page for it serves the identical abstract with the same three figures.

Personalization won on the metric a recommender is usually judged by. It lost on the one underneath it. Both numbers came out of the same experiment.

  • Design: users were randomized between personalized podcast recommendations and recommendations of the most popular podcasts among users in the same demographic group. The comparison is causal, not a before-and-after on logs.
  • Short-term engagement: the treatment raised podcast streams by 28.90% (+/- 3.81%). That is a large, real win on the metric a recommender is usually judged by.
  • Individual diversity: in the same experiment, the average individual-level diversity of podcast streams fell by 11.51% (+/- 1.08%).
  • Aggregate diversity: across users, diversity of streams rose by 5.96%. The catalog-level number moved in the opposite direction from the per-user number, so which one you report decides what you conclude.
  • Interpretation: the authors read the pair as “patterns of consumption that are homogenous within and diverse across users, a pattern reflecting Balkanization”. A single aggregate diversity metric would have scored that as an improvement.

Visual

A recommendation scorecard, part of it written by law

Quality spans users, items, and the ecosystem. Since 2022, part of this scorecard is no longer the team's to choose. The EU's Digital Services Act requires online platforms to set out “the main parameters used in their recommender systems” in plain and intelligible language, including the criteria most significant in determining what is suggested. That is Article 27. Where there is an option to switch, it has to be “directly and easily accessible”. Article 38 goes further for the largest services: “In addition to the requirements set out in Article 27, providers of very large online platforms and of very large online search engines that use recommender systems shall provide at least one option for each of their recommender systems which is not based on profiling as defined in Article 4, point (4), of Regulation (EU) 2016/679.”

Read against the rows below, that turns coverage into an external specification. At least one reachable ranking that is not built on profiling, plus a published account of what drives the one that is. Ireland's Digital Services Coordinator, Coimisiun na Mean, states the same obligation in user terms — “a right to choose a recommender system feed that is not based on the profiling of their personal data”, with the switch available at any time and easily accessible.

The long-term health row now has enforcement attached to it. On 5 May 2026 Coimisiun na Mean opened two separate formal DSA investigations into Meta, one over Facebook and one over Instagram. They assess contravention of Article 27 (in particular 27.3) and Article 25 (in particular 25.1). Can users select and modify a recommender feed not based on profiling, through a directly and easily accessible functionality? And do the interfaces deceive or manipulate users away from choosing it? A finding of non-compliance can carry an administrative financial sanction of up to 6% of turnover.

The regulator did not frame this as a technical matter. John Evans, Digital Services Commissioner at Coimisiun na Mean, put it this way: “Our message is clear: it is unacceptable for platforms to prevent people from using their rights under the law, or to try to manipulate people away from making empowered choices about whether or not recommender system feeds control what they see online.”

What is being measured there is not relevance at all. It is whether an alternative ranking is reachable.

FigureHierarchy · 5 levels
  • Relevance

    Do recommended items match the user’s current need or preference?

    • Coverage

      How much of the user base and catalog receives meaningful service or exposure?

      • Diversity and novelty

        Does the list avoid redundancy and surface unfamiliar but plausible options?

        • Calibration

          Does the recommendation mix reflect the user’s demonstrated interests without becoming repetitive?

          • Long-term health

            What happens to satisfaction, retention, creators, inventory, and feedback loops?

Key idea

Logged-policy evaluation needs assumptions

Offline replay can evaluate only actions with adequate support under the logging policy. If a new policy recommends items that were rarely or never shown, logged outcomes provide little direct evidence. Propensity weighting and doubly robust methods can help under explicit assumptions. Extreme weights, hidden confounding and incorrect logging remain serious risks.

The repair was formalised in 2016, and the idea is in the name: treat recommendations as treatments. Schnabel, Joachims and colleagues handle the selection biases in recommendation data by “adapting models and estimation techniques from causal inference”, an approach that “leads to unbiased performance estimators despite biased data”.

The guarantee rests on the propensity model. A larger log does not supply one. That is how an estimator can be unbiased in principle and unusable on exactly the actions a new policy favours.

No estimator can recover outcomes for actions that the data never meaningfully explored without additional assumptions.

Users and items have cold-start regimes

Established users with rich histories often dominate offline metrics. New users, new items, sparse-interest users and long-tail creators can experience a different system.

Report warm- and cold-start results, exposure concentration, creator or supplier slices, and fallback quality. A global ranking metric can hide entire populations with weak support.

A 2019 study showed how much it can hide. Abdollahpouri and colleagues split users by their interest in popular items into three groups they called “Niche, Diverse and Blockbuster-focused”. Recommendations came back “extremely concentrated on popular items even if a user is interested in long-tail and non-popular items showing an extreme bias disparity”. The niche users were served the blockbusters anyway. A catalog-coverage number computed over the whole user base would have averaged that away and shown nothing.

Recommendation quality should cover both sides of the marketplace when both are affected.

Analogy

A bookstore that learns only from its front table

Measuring demand only for the books on the front table, then using those sales to choose next week's table, is a bookstore reading its own handwriting. Popularity partly reflects exposure created by the store itself.

The store cannot know what would have sold from a shelf it never used. Next week's table is chosen out of exactly that gap. What a recommender shows is what it will later be scored on.

The 2005 eye-tracking experiment is that bookstore under laboratory control. The proxy moved the books between the table and the back shelf without changing a single title, and the sales followed the furniture — mean clicked rank 2.66 to 4.03. A store that never runs that manipulation has no way to tell the two explanations apart.

Exposure is both an outcome and a source of future labels.

Steps

Move from offline ranking to ecosystem evidence

Use staged evaluation, because no one dataset answers every question.

Step 1 is where the Netflix Prize stopped, and it is worth knowing how that ended. On 26 July 2009 BellKor's Pragmatic Chaos submitted an entry scoring RMSE 0.8567 on the test set, a 10.06% improvement on Netflix's Cinematch baseline. Twenty minutes later The Ensemble submitted an entry that also scored 0.8567. To two further decimal places the test-set scores were 0.856704 for the winners and 0.856714 for the runners-up. Feuerverger and colleagues, writing in Statistical Science in 2012, put it plainly: “Since the contest rules were based on test set RMSE, and also were limited to four decimal places, these two submissions were in fact a tie. It is therefore the order in which these submissions were received that determined the winner”. The National Institute of Statistical Sciences reported the same 0.8567 and the same 10.06% when the award was made on 21 September 2009.

A $1M prize was settled by the clock. The single offline accuracy number had run out of resolution while two very different systems were still indistinguishable on it. Steps 2 through 5 exist for that reason: portfolio diagnostics, counterfactual support checks, an online experiment with guardrails, and a delayed review of retention, content health, creator exposure and feedback loops. An accuracy figure can be the best one available and still unable to tell two candidates apart.

FigureProcess · 5 steps
  1. 1. Offline relevance

    Measure top-k ranking on a leakage-safe temporal split.

  2. 2. Portfolio diagnostics

    Add coverage, diversity, novelty, calibration, and concentration.

  3. 3. Counterfactual checks

    Audit logging support and estimator stability for policy changes.

  4. 4. Online experiment

    Randomize eligible users with satisfaction and safety guardrails.

  5. 5. Long-term review

    Track delayed retention, content health, creator exposure, and feedback loops.

Key takeaways