Research
Do LLMs Benefit from User and Item Embeddings in Recommendation Tasks?
Overview Research area: LLM-based recommender systems, specifically the integration of collaborative filtering signals into generative language models. Technical level: Intermediate. It assumes famili

- arXiv
- 2601.04690
- Published
- 2026-01-08
- Authors
- Mir Rayat Imtiaz Hossain, Leo Feng, Leonid Sigal, Mohamed Osama Ahmed
AI summary
Overview
- Research area: LLM-based recommender systems, specifically the integration of collaborative filtering signals into generative language models.
- Technical level: Intermediate. It assumes familiarity with collaborative filtering, embeddings, LoRA fine-tuning, and standard top-k recommendation metrics, but the architecture itself is described simply.
- Scope: The paper proposes and empirically evaluates a framework that projects user and item embeddings learned via matrix factorization into an LLM's token space through separate lightweight projectors, comparing it against text-only LLM baselines and traditional recommender systems on three benchmark datasets.
What This Paper Is About
LLMs used as recommenders typically work from text alone, which means they cannot directly capture the co-occurrence and interaction patterns that collaborative filtering methods exploit. Some prior work injects user or item embeddings, but these approaches are limited to binary classification, handle only a single target item embedding at a time, or use item embeddings while ignoring user embeddings. This paper's goal is to let a finetuned LLM condition simultaneously on one user embedding and an arbitrary number of item embeddings from the user's history, alongside text tokens, so that richer collaborative structure informs the recommendation.
Key Contributions
- A framework that injects both user and item collaborative filtering embeddings into an LLM, allowing conditioning on a single user embedding plus an arbitrary number of item embeddings from the interaction history.
- Two separate lightweight projectors (user projector and item projector), each a two-layer MLP, that independently map collaborative embeddings into the language space so the model can learn complementary representations.
- A two-stage fine-tuning strategy: Stage 1 freezes the LLM and trains only the projectors; Stage 2 jointly fine-tunes the projectors with LoRA adapters on the LLM backbone.
- An empirical comparison against text-only LLM baselines (OpenP5's Llama-R, Llama-S, Llama-C) and traditional recommender systems on three OpenP5 datasets across both Sequential and Straightforward recommendation tasks.
Main Findings
- Text-only LLMs perform poorly: Without collaborative embeddings, decoder-only LLM baselines (Llama-R, Llama-S, Llama-C) show very low scores. For example, on Amazon Beauty in the sequential task, Llama-C reaches HR@5 of 0.0002 and Llama-R reaches 0.0018.
- Stage 2 beats Stage 1 consistently: Moving from projector-only training to joint optimization of projectors and LoRA adapters produces substantial gains. On Amazon Beauty sequential, HR@5 rises from 0.0361 (Stage-1) to 0.0642 (Stage-2).
- Strong sequential results that surpass classical baselines on Amazon Beauty: The Stage-2 model reaches HR@5 0.0642, NDCG@5 0.0514, HR@10 0.0794, and NDCG@10 0.0563 on Amazon Beauty, exceeding SASRec (HR@5 0.0387, NDCG@5 0.0249, HR@10 0.0605, NDCG@10 0.0318) and HGN (HR@5 0.0325, NDCG@5 0.0206, HR@10 0.0512, NDCG@10 0.0266).
- User-embedding-only setting is more robust than traditional systems: In the straightforward task, where only the user ID is given, the method retains strong performance on 2 of 3 datasets (Beauty and LastFM) and outperforms traditional recommender systems. On Beauty, Stage-2 achieves HR@5 0.0586 versus SimpleX at 0.0300, BPR-MF at 0.0224, and BPR-MLP at 0.0193. On LastFM, Stage-2 achieves HR@5 0.0505 versus SimpleX at 0.0312. On MovieLens-1M the method scores HR@5 0.0205, HR@10 0.0419, NDCG@10 0.0185, while SimpleX reaches HR@5 0.0301 and HR@10 0.0596.
- Robustness to unseen prompts: Under unseen prompt templates, text-only Llama variants degrade, while the embedding-conditioned model maintains strong performance. On Amazon Beauty sequential unseen prompts, Stage-2 records HR@5 0.0631 and NDCG@5 0.0507, close to its seen-prompt score of HR@5 0.0642.
- ILM comparison caveat: ILM results are reported for Amazon Beauty and MovieLens-1M in both tasks, but are marked as not reported ("-") for LastFM.
Methodology in Plain English
The researchers take a text prompt describing a user and their interaction history and pass it through an LLM tokenizer. Instead of relying on the model to interpret the user ID and item IDs as text, they look up precomputed collaborative filtering embeddings for those IDs. These embeddings come from matrix factorization on the user-item interaction data using Weighted Alternating Least Squares (WALS), where the dot product between a user embedding and an item embedding approximates the preference score. The paper notes that more sophisticated collaborative filtering techniques could be applied, and that matrix factorization was chosen as a simple, effective foundation.
Because the LLM expects representations in its own language space, two separate two-layer MLP projectors convert the user embedding and the item embedding into token-space embeddings, which are then injected alongside the ordinary text tokens. Training uses standard next-token prediction loss. The authors first train only the projectors with the LLM frozen (Stage 1), then jointly fine-tune the projectors and LoRA adapters on the LLM backbone (Stage 2).
Experiments use three OpenP5 datasets (Amazon Beauty, LastFM, MovieLens-1M), each with Sequential and Straightforward recommendation tasks. For prompts, they follow OpenP5's 11 templates per task, training on 10 and reserving 1 for zero-shot generalization. The backbone is OpenLLaMA-3B, a reproduction of LLaMA-2, and training alternates between the two tasks. Evaluation uses top-k Hit Ratio and NDCG at k=5 and k=10.
Why This Matters
The work shows a practical route for connecting traditional collaborative filtering, which excels at capturing user-item structure, with generative LLM recommenders, which excel at language understanding and generalization. It also demonstrates that text-only LLM recommenders are weak on these benchmarks, and that grounding them in structured interaction signals both improves accuracy and improves robustness to prompt variation.
Real-world applications:
- E-commerce product recommendations: Using purchase and browsing history to suggest the next item, where item co-occurrence patterns matter more than product text.
- Music and media streaming: Next-track or next-title prediction based on listening history, as tested on LastFM.
- Content and video platforms: Sequential next-item prediction for users with long interaction histories.
- Personalized marketing or financial product suggestions: Where only user-level signals may be available at inference time, matching the straightforward setting.
Industry relevance: the framework adds only small projector modules and LoRA adapters on top of an existing LLM, rather than training a large model from scratch, which is attractive for organizations that already operate collaborative filtering pipelines and want to layer generative recommendation on top. The industrial author affiliations (RBC Borealis, University of British Columbia) suggest direct interest from applied research settings.
Future Directions
- Encoding richer item representations, such as semantic embeddings derived from product descriptions, rather than relying only on ID-based matrix factorization embeddings.
- Exploring more advanced recommendation system architectures beyond the matrix factorization setup used here.
- Developing improved methods for encoding user-item interactions within the LLM.
- Determining whether the robustness advantage on unseen prompts and the user-embedding-only setting can be extended to the dataset where the method underperformed relative to traditional systems (MovieLens-1M in the straightforward task).
Target Audience
Researchers and practitioners working at the intersection of LLMs and recommender systems, particularly those interested in multimodal-style injection of non-text embeddings into language models. It is also relevant to applied machine learning engineers who maintain collaborative filtering systems and want to augment them with generative models, and to students who need a concise example of a projector-based architecture combined with a two-stage fine-tuning recipe.
Note: The provided paper content does not report dataset sizes for Amazon Beauty, LastFM, or MovieLens-1M.
Authors’ abstract
Large Language Models (LLMs) have emerged as promising recommendation systems, offering novel ways to model user preferences through generative approaches. However, many existing methods often rely solely on text semantics or incorporate collaborative signals in a limited manner, typically using only user or item embeddings. These methods struggle to handle multiple item embeddings representing user history, reverting to textual semantics and neglecting richer collaborative information. In this work, we propose a simple yet effective solution that projects user and item embeddings, learned from collaborative filtering, into the LLM token space via separate lightweight projector modules. A finetuned LLM then conditions on these projected embeddings alongside textual tokens to generate recommendations. Preliminary results show that this design effectively leverages structured user-item interaction data, improves recommendation performance over text-only LLM baselines, and offers a practical path for bridging traditional recommendation systems with modern LLMs.