Skip to content
AI.info

Research

Personal-Agent Mediated Recommendation with Cross-Platform User History

Overview Research area: Recommender systems, LLM-based personal agents, and reinforcement learning from recommendation feedback. Technical level: Advanced. The paper combines a new task formalization,

Personal-Agent Mediated Recommendation with Cross-Platform User History
arXiv
2610.07588
Published
2026-10-06
Authors
Yu Xia, Jiangfan Zhang, Jun Xiao, Julian McAuley, Xiangjun Fan

AI summary

Overview

  • Research area: Recommender systems, LLM-based personal agents, and reinforcement learning from recommendation feedback.
  • Technical level: Advanced. The paper combines a new task formalization, a benchmark construction, and a theoretically analyzed policy-optimization method involving rank-aware advantage decomposition, KL-regularized allocation, and projection onto a constraint cone.
  • Scope: The paper formalizes "Personal-Agent Mediated Recommendation," introduces the MediateRec benchmark (four Amazon-based proxy platforms plus a real cross-platform OpenPlay test), and proposes Personal Attribution Mediation Optimization (PAMO) for training a personal agent to selectively revise a strong platform ranking using cross-platform history.

What This Paper Is About

Recommender systems today personalize within a single platform, but users increasingly could have a personal LLM agent that holds their history across many services and acts on their behalf. The paper asks how such an agent should mediate an existing platform ranking: when does cross-platform history justify overriding a strong platform recommendation, and when should the agent defer to it? The goal is to define this setting, build a benchmark for it under a controlled platform–agent information boundary, and train an agent that achieves a better balance between beneficial "rescues" and harmful overrides.

Key Contributions

  1. Formalization of Personal-Agent Mediated Recommendation. The paper defines the setting where a platform recommender ranks a candidate set using platform-local information, and a personal agent receives that proposal plus user-authorized cross-platform history to produce the final top-K slate. It characterizes the resulting platform-relative outcomes (preserved hits, rescues, harmful overrides, mutual misses) and the rescue–harm trade-off that distinguishes mediation from standalone reranking.

  2. The MediateRec benchmark. MediateRec includes scalable proxy cross-platform environments built from Amazon Reviews, covering target categories Movies & TV (Movie), Toys & Games (Toy), Grocery & Gourmet Food (Grocery), and Beauty & Personal Care (Beauty), plus a real cross-platform external test built from OpenPlay, which links Steam, Nintendo, and Xbox activity through a shared pseudonymized person identifier. The benchmark enforces a controlled platform–agent information boundary: the platform proposal is generated from platform-local information, and cross-platform history is exposed only to the personal agent. The paper states MediateRec will be released upon internal approval.

  3. Personal Attribution Mediation Optimization (PAMO). PAMO counterfactually masks cross-platform history to estimate a "personal mediation support" score for each sampled rationale, decomposes NDCG into rank-aware cutoff-level advantages, and reallocates advantage mass toward more strongly supported platform-disagreeing responses under a platform-relative value floor. The paper proves that PAMO preserves cutoff-level positive and negative advantage mass and is locally optimal among first-order reallocations that preserve this mass without lowering average platform-relative value.

Main Findings

  • Platform defaults are strong. The frozen platform ranking (a target-platform-trained SASRec model followed by Claude Sonnet 4.6 reranking using only within-platform history) achieves mean HR@10 / NDCG@10 of 49.3% / 0.329 across the four Amazon datasets and 40.8% / 0.241 on OpenPlay.

  • Cross-platform history alone is not enough. Results on MediateRec show that access to cross-platform history does not by itself ensure effective mediation: strong proprietary LLMs make meaningful platform corrections yet still introduce non-negligible harmful overrides.

  • PAMO improves over matched outcome-only RL. PAMO consistently improves over matched outcome-only RL (GRPO with NDCG@10 reward) across seen target platforms, unseen target platforms, and the real cross-platform test, with better recommendation accuracy and a better rescue–harm balance. The exact numerical values of these improvements are not reported in the available paper content.

  • Theoretical guarantees hold at the cutoff level. Proposition 1 establishes that PAMO preserves aggregate positive and negative group-relative advantage mass at each active cutoff and does not decrease the average direction-aligned NDCG magnitude or average personal mediation support relative to uniform allocation.

  • PAMO's reallocation direction is locally optimal. Theorem 1 shows that the right derivative of the PAMO weights at the uniform allocation is the Euclidean projection of the centered support score onto the cone of admissible first-order reallocations, and that this direction maximizes the first-order increase in average personal mediation support among unit-norm admissible reallocations.

  • Scope of the theory is limited. The paper states that these results characterize advantage allocation within a fixed sampled rollout group, treating the support score and value magnitude as fixed, and do not imply global optimality of the neural policy or a guaranteed held-out utility improvement.

  • Measured quantities. Evaluation reports HR@{3,5,10} and NDCG@{3,5,10}, plus rescue, harmful override, and intervention, where intervention is the number of platform-slate items replaced in the final slate, d(S, P) = K − |S ∩ P|.

Methodology in Plain English

The setup is a single recommendation episode per user. The platform produces a ranked proposal over 50 candidates: the target item plus 49 items the user has not previously interacted with, sampled in proportion to item popularity. For Amazon, the platform is built from within-category history plus population interactions via SASRec, then reranked by Claude Sonnet 4.6 at temperature 0 using only within-platform history. That final ranking is frozen for everyone. The personal agent sees the platform proposal, within-platform history, and cross-platform history, but not platform scores, population interactions, or model states, and returns a ranked top-10 in a single call.

For training, PAMO starts with a Qwen3-4B-Instruct-2507 backbone initialized on 2,000 Claude Sonnet demonstrations (1,000 each from Movie and Toy, not filtered by outcome). It then samples 8 rollouts per prompt. For each rollout it computes a support score by taking the generated rationale and comparing its likelihood under the full history versus a version where cross-platform history has been masked out; the rationale is scored rather than the final item identifiers because the identifiers can become less sensitive to the removed history once a fixed rationale prefix is teacher-forced.

Outcome supervision comes from decomposing NDCG@K into nested top-k cutoffs, each with a marginal utility, and then computing group-relative advantages at each cutoff as in GRPO. Only rollouts that disagree with the platform at a given cutoff are reallocated: rescues when the platform missed, harmful overrides when the platform succeeded. Within that disagreement set, PAMO solves a constrained optimization that tilts advantage toward higher-support responses while requiring the average direction-aligned NDCG magnitude to stay at least as high as the uniform allocation. The solution is an exponential-family weighting with a dual parameter on the value constraint, and the strength of the tilt is controlled by a concentration parameter κ with the inverse temperature capped at η_max = 60. Cutoff advantages are recombined using the same marginal utilities that define NDCG, then used in the same clipped token-level GRPO objective as the baseline.

Why This Matters

  • Research impact: The paper shifts attention from platform-side agentic recommenders to a user-side personal agent that mediates an existing, strong platform proposal, and it provides a benchmark with an explicit platform–agent information boundary plus a real cross-platform external test. It also contributes a credit-assignment idea—masking cross-platform history to attribute advantage—that is distinct from standard outcome-only RL for recommendation.

  • Real-world applications:

    • Consumer assistants that hold a user's purchase, viewing, or listening history across several services and can adjust a storefront's recommendations on the user's behalf.
    • Cross-service content or product discovery, where an agent knows a user's behavior on one platform and can correct misses on another.
    • Gaming storefronts, motivated directly by the OpenPlay setup linking Steam, Nintendo, and Xbox activity under a shared pseudonymized identifier.
    • Privacy-conscious personalization, since the design keeps cross-platform history on the user side and never exposes it to the platform during ranking construction.
  • Industry relevance: The setting maps onto an architecture in which platforms keep their population-level collaborative models and users govern an agent over their own cross-service data. The paper's finding that even strong proprietary LLMs produce non-negligible harmful overrides suggests that deploying such mediators requires explicit control of the rescue–harm trade-off rather than relying on a capable base model, and PAMO adds only one masked teacher-forced pass during training with no extra rollouts and no inference-time component.

Future Directions

  • Determine when cross-platform evidence should override a platform decision. The central open question the paper poses—when does cross-platform history justify changing a strong platform ranking, and when should the agent defer—remains a learning problem the authors only partially address.
  • Extend beyond a frozen platform proposal. The setting holds the target platform recommender fixed; how mediation interacts with a platform that adapts to agent interventions is not addressed.
  • Improve the theory beyond local optimality. The guarantees are first-order and within a fixed sampled rollout group, so whether PAMO's reallocation leads to global policy improvements or guaranteed held-out utility gains is left open.
  • Broaden benchmark coverage. MediateRec's proxy environments are Amazon categories rather than independent services, and the real cross-platform test is a single 645-user OpenPlay cohort; scaling real cross-platform evaluation and adding more linked services are natural next steps.

Target Audience

Researchers and practitioners working on recommender systems, LLM agents, and reinforcement learning for user-facing policies—particularly those interested in user-governed personalization, cross-domain or cross-platform recommendation, and credit assignment methods that go beyond outcome-only rewards. Readers need comfort with policy-gradient methods, group-relative advantage estimation, and constrained optimization to follow the theory sections.

Authors’ abstract

Modern recommendation is shifting from platform-centric personalization toward user-governed personalization, where a personal LLM agent can act on the user's behalf across services. We formalize this emerging paradigm as Personal-Agent Mediated Recommendation: a platform recommender ranks a candidate set using platform-local information, and a personal agent uses user-authorized cross-platform history to mediate the resulting ranking and produce the final top-K slate. Such mediation is nontrivial: the platform ranking can encode strong population evidence that the personal agent cannot observe, so effective mediation must therefore balance beneficial rescues against harmful overrides. To study this trade-off, we introduce MediateRec, a benchmark that includes scalable proxy cross-platform environments and a real cross-platform test under a controlled platform-agent information boundary. To train the agent to use cross-platform history effectively, we further propose Personal Attribution Mediation Optimization (PAMO), which counterfactually masks that history to estimate personal mediation support and reallocates rank-aware advantage mass under a platform-relative value floor. We theoretically prove that PAMO preserves cutoff-level advantage mass and is locally optimal among first-order reallocations that preserve this mass without lowering average platform-relative value. Experiments on MediateRec show that personal-agent mediation enables meaningful platform corrections, yet even strong proprietary LLMs introduce non-negligible harmful overrides. PAMO consistently improves over matched outcome-only RL across seen and unseen target platforms and on the real cross-platform test, while achieving a better rescue-harm balance.

Read the original paper