Research
RPTune: Learned Context Curation for LLM Catalog Search
RPTune: Learned Context Curation for LLM Catalog Search Overview Research area: Information retrieval (cs.IR), specifically e-commerce product search — combining long-context LLM prompting, dense retr

- arXiv
- 2610.00964
- Published
- 2026-10-01
- Authors
- Chuxuan Hu, Hejie Cui, Norman Huang, Shubham Kumar Bharti, Wang-Chiew Tan, Sercan Ö. Arık
AI summary
RPTune: Learned Context Curation for LLM Catalog SearchOverview
Research area: Information retrieval (cs.IR), specifically e-commerce product search — combining long-context LLM prompting, dense retrieval/ranking, and LLM post-training (reinforcement learning).
Technical level: Intermediate. The paper assumes familiarity with retrieval-augmented generation, dense encoders, and GRPO-style reinforcement learning, but its central argument (fit the whole catalog in the prompt, then curate and adapt) is explained plainly.
Scope in one sentence: The paper proposes a framework that learns how to prune and order a small merchant's product catalog before feeding it to an LLM, then post-trains the LLM on those curated catalogs, evaluated over 700 complex conversational queries (100 each from 7 real merchants).
What This Paper Is About
Existing product search systems are built for large marketplaces with millions of items, dedicated ML teams, and rich behavioral logs — assumptions that do not hold for small merchant businesses (SMBs). The paper studies in-context catalog search: for the 56 real-world Shopify storefronts the authors analyze, 92.9% of catalogs fit within 1M tokens, with a median catalog size of only 51K tokens, so an LLM can in principle reason over the entire catalog at once instead of running a multi-stage retrieval pipeline.
The problem is that fitting the catalog into the context does not mean the model uses it well — LLMs do not use long contexts uniformly. The goal is therefore to answer two questions: (1) how to curate and present catalogs to the LLM, and (2) how to adapt the LLM for product selection on those curated contexts.
Key Contributions
-
A framework (RPTune) coupling learned context curation with LLM post-training. RPTune combines an encoder that embeds the query and products with a reorganizer that selects and positions products, feeding a compact, query-dependent context to a downstream LLM. The name stands for "Rank, Prune, and Fine tune."
-
An LLM-in-the-loop reorganizer trained with reinforcement learning. Rather than fitting the reorganizer to relevance labels directly, the authors sample selections from a frozen downstream LLM, score them against catalog-wide relevance scores, normalize rewards within each query group (GRPO-style), and use the advantages to update only the reorganizer's MLP parameters while the encoder and LLM stay frozen.
-
A context-relative post-training reward. The LLM is post-trained with GRPO using a reward of +1 when it selects the highest-relevance product available in the curated context, −1 when it outputs a product outside the curated catalog (penalizing hallucination), and 0 otherwise — so a positive learning signal survives even when curation omits the globally best variant.
-
Fully synthetic, catalog-grounded supervision requiring no merchant relevance labels or user logs, plus an evaluation over 7 real merchants spanning distinct retail verticals with 100 complex conversational queries per merchant (700 total).
Main Findings
-
Full-catalog prompting beats retrieval pipelines (RQ1). With gemini-3.7-flash as a shared backbone, directly prompting the LLM with the full catalog outperforms all baselines by at least 9.0 EM points, while maintaining single-digit latency (8 s). Separately, full-catalog LLM search improves accuracy over prior retrieval-based approaches by up to 20 percentage points.
-
Curation helps across backbones (RQ2). Across 9 frozen LLM backbones, RPTune improves exact-match accuracy by 13.5 percentage points and feature reward by 10.1 points on average, with maximum EM gains reaching 20 points. Gains appear in all 63 backbone–merchant combinations. The paper also reports that across 8 evaluated frontier LLMs, curation gains average 14.0 percentage points and reach up to 31.4 points.
-
The gains come from the learned curator, not generic pruning. On gemini-3.7-flash, substituting Jina Reranker 3.5 or Gemini Embedding 2 at the same 25% retention budget degrades accuracy on 5 and 3 of the 7 merchants respectively, whereas RPTune improves all 7.
-
Curation also reduces latency. Despite the extra curation step, RPTune reduces end-to-end latency for all nine LLMs by 28% on average and by up to 71%.
-
Post-training adds further gains (RQ3). Combining curation with post-training nearly triples EM over the raw gemma-4-E4B-it backbone (from 10.7% to 31.0%), an additional 10.3 EM and 9.3 FR points beyond inference-time curation alone. Performance increases monotonically across all seven merchants through the stages raw backbone → encoder-only → full curator → post-training.
-
Case study on Beauty Bakerie. For the example bridal query, RPTune places the ground-truth product Cherry Flambé Matte Lip Whip (15 g) at position 8 of a compact context, raising EM from 0% to 100% for grok-4.1-fast-reasoning and from 0% to 70% for gemma-4-E4B-it; post-training then raises gemma's EM from 70% to 100%. Before post-training, gemma sometimes picked 23 g alternatives as "close to the lightweight preference" even though the 15 g ground truth was available.
-
RPTune generalizes to new products without retraining. Trained on a reduced Beauty Bakerie catalog (half of the products removed within each product type), post-trained RPTune shows only a 14.7% relative EM decline as the catalog doubles, versus a 29.5% drop for the baseline, with FR practically unchanged. At double the initial catalog size it retains absolute gains of 35.0 EM and 19.2 FR over the baseline.
-
RPTune generalizes to unseen merchants. All 7 × 7 train/eval merchant combinations improve both EM and FR over the gemma baseline, with average out-of-domain gains of +18.1 EM and +13.6 FR, versus in-domain +20.3 and +17.7, despite a mean pairwise Jaccard similarity of 0.146 across catalogs.
-
Ablations on the post-training objective and curation. Curation improves every tested post-training method (DPO, IRPO, GRPO), averaging +22.2 pp EM and a 34% latency reduction; GRPO performs best and benefits most from RPTune. RPTune's context-relative reward matters — on Beauty Bakerie, EM drops 10 points when replaced with a continuous reward or a global top-1 reward. Training on curated contexts improves EM by 5.7 points when evaluating on raw contexts and 9.0 points when evaluating on curated contexts; curation at both training and inference cuts latency by 35% versus the raw baseline.
-
Pruning and ordering sweet spots. Sweeping retention from 5% to 100%, RPTune beats the raw full-catalog baseline at every budget, including full retention. The default 25% retention yields 77.1% average context recall (over 3× better than random selection), and 50% retention reaches 90.7%. On Beauty Bakerie, EM and FR peak at 25% retention (with 92% recall). Every RPTune curated layout beats all baseline layouts by 16.8 EM points while cutting latency by 32%, with ascending order best, followed by U-shaped.
-
LLM feedback is what makes the reorganizer work. Replacing the RL objective with a supervised listwise cross-entropy fitted to the relevance labels yields 16.4% average EM with frozen gemma — no better than the encoder alone and 4.3 points below RPTune's reorganizer, which wins on all 7 merchants.
-
Cross-model efficiency. RPTune context curation delivers accuracy gains comparable to a model upgrade with up to 21× speedup, and post-training brings gemma-4-E4B-it to accuracy levels comparable to frontier models.
Methodology in Plain English
Setup. Each merchant's catalog is a set of product variants; each query is a single-turn natural-language request that mixes hard requirements with soft preferences. The system must output the best-matching variant. The catalog is the only merchant-provided data source — no relevance labels, no user behavior logs.
Inference. RPTune has two parts. An encoder turns the query and every product into L2-normalized dense vectors. A reorganizer then scores each product with a combination of base embedding similarity (scaled by τ = 10) and a learned MLP adjustment over the concatenated query–product vector. It keeps the top k = ⌈(1−ρ)n⌉ products for a pruning rate ρ (25% retention by default) and arranges them in ascending priority order, so the highest-scoring products sit closest to the generation prompt. The LLM then picks a product from this compact, curated context.
Training data. Everything is synthetic and grounded in the merchant's own catalog. For each product variant, gemini-3.1-pro generates 10 diverse 1–2 sentence queries targeting it (10n queries total). Then the same LLM acts as a relevance judge, assigning a graded score from 0.0 to 1.0 to every catalog variant for each query, scoring in batches of 10 variants per call with a running dictionary of prior scores so the ranking stays globally consistent. Note that the original target product is not guaranteed a score of 1.0.
Three training stages, each updating one module while the rest stay frozen.
- Encoder: trained with Multiple Negatives Ranking Loss on the query–product pairs, pulling matching pairs together relative to in-batch negatives.
- Reorganizer: trained with LLM-in-the-loop reinforcement learning. For each query, the frozen LLM (gemma-4-E4B-it) produces G selections at non-zero temperature from the same curated context; rewards come from the relevance table, with −1 for outputs outside the catalog. Rewards are standardized within each query group to get advantages (GRPO-style), and a softmax distribution over product scores serves as a differentiable link to the reorganizer's parameters. Only the MLP is updated; its output layer is zero-initialized so training starts from the encoder's ordering.
- LLM post-training: with the curator frozen, GRPO trains the LLM on curated contexts using the context-relative reward (+1 for the best product available in the curated context, −1 for hallucinating a product outside it, 0 otherwise).
Evaluation. Seven real Shopify merchants spanning different retail verticals, each with 100 synthetic multi-constraint conversational queries built by a three-stage pipeline: gemini-3.1-pro generates queries from high-level merchant context only (withholding individual product records), selection consensus across randomized catalog orderings retains moderate-difficulty queries (31.9% of candidates), and human experts verify pilot queries and audit the final set (9.6% rejection rate). Metrics are exact match (EM), feature reward (FR), and end-to-end latency; each method runs 10 times on the same query set and results are averaged over queries and runs. Baselines include RAG (top-10), RAG-Fusion, TourRank, LongLLMLingua, dense retrieval with EmbeddingGemma 300M and Gemini Embedding 2, and listwise ranking with Jina Reranker 3.5, plus DPO and IRPO for the post-training comparison.
Why This Matters
The paper reframes a design assumption: for small merchants, the right architecture may not be a miniature version of a large-marketplace retrieval stack, but rather a curated full-catalog prompt. It shows that curation quality — not just whether the catalog fits — is the dominant lever, and that the LLM's own behavior is a useful training signal for deciding what to put in the context. It also shows that both curation and post-training can be learned from synthetic, catalog-grounded data, removing the dependency on merchant relevance labels or user logs.
Real-world applications:
- SMB e-commerce search on Shopify-class storefronts, where catalogs are small enough to fit in context and merchant ML teams do not exist.
- Conversational shopping assistants handling multi-constraint, trade-off-heavy queries (color, finish, weight, price, occasion) rather than keyword lookups.
- Inventory-update resilience: merchants constantly change products, and RPTune transfers to new products and unseen merchants without retraining.
- Cost and latency reduction for LLM-powered search, since proactive pruning shortens prompts and cuts end-to-end latency (28% on average, up to 71% across the nine backbones tested).
Industry relevance: the results target the economics of running product search for long-tail merchants — the paper's framing is that curation can substitute for a model upgrade (up to 21× speedup at comparable accuracy), and that post-training can bring open-weight models such as gemma-4-E4B-it to accuracy levels comparable to frontier models.
Future Directions
- Relaxing the long-context assumption. The method applies only to catalogs that fit in the LLM's context window alongside the query and instructions; how the curator should behave for catalogs that do not fit is not reported.
- Better selection of the pruning budget and ordering. Downstream accuracy peaks at an intermediate budget (25% on Beauty Bakerie, with 92% recall) while larger budgets add recall but lower accuracy and raise latency; a principled way to pick ρ and the layout per merchant is left open.
- Extending beyond synthetic supervision. All training uses LLM-generated queries and LLM-judged relevance; the paper reports a 9.6% human rejection rate for evaluation data, but how far the training signal can be pushed without any real behavioral data remains an open question.
- Broader generalization testing. Zero-shot transfer is demonstrated across 7 merchants and 7 × 7 merchant pairs with a mean pairwise Jaccard similarity of 0.146; whether the same holds across larger, more heterogeneous catalog families is not established here.
Target Audience
Practitioners and researchers in information retrieval and e-commerce search who work with LLMs; engineers building search or shopping assistants for small or mid-sized merchants; and researchers interested in context curation, long-context prompting, and reinforcement-learning post-training with automatically generated supervision. Readers should be comfortable with retrieval pipelines, dense embeddings, and group-relative policy optimization, though the paper's framing (curate first, then adapt the model) is accessible without deep prior knowledge of those areas.
Authors’ abstract
For small merchant businesses (SMBs) whose catalogs fit within a long-context LLM, full-catalog prompting offers a compelling alternative to multi-stage retrieval designed primarily for large marketplaces with millions of items. However, fitting the full catalog into the context window does not ensure that the model can use it effectively, since LLMs do not exploit long contexts uniformly. We therefore study in-context catalog search through two complementary questions: (1) how to curate and present catalogs to the LLM, and (2) how to adapt the LLM for product selection on curated contexts. We propose RPTune, an end-to-end framework that couples learned catalog curation with LLM post-training using automatically generated, catalog-grounded supervision. An encoder-reorganizer curator orders and prunes products guided by downstream LLM feedback, while the resulting curated catalogs in turn improve the effectiveness of LLM post-training with a context-relative reward. We evaluate RPTune on 7 real merchants spanning distinct retail verticals, using 100 complex conversational queries per merchant. RPTune consistently improves search accuracy across both proprietary and open-weight LLMs, with context curation yielding gains of up to 31.4 percentage points and post-training adding a further 10.3 points on average.