Research
It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning
CoGR: Co-Evolving Generative Retriever with Reinforcement Learning Overview Research area: Information retrieval (cs.IR) — specifically generative/lexical retrieval, LLM-based query and item represent

- arXiv
- 2609.00638
- Published
- 2026-09-01
- Authors
- Runpeng Dai, Kaili Huang, Changsung Kang, Ciya Liao
AI summary
CoGR: Co-Evolving Generative Retriever with Reinforcement LearningOverview
Research area: Information retrieval (cs.IR) — specifically generative/lexical retrieval, LLM-based query and item representation learning, and reinforcement learning for retrieval optimization.
Technical level: Intermediate (readers should be comfortable with retrieval metrics, supervised fine-tuning, and the basics of policy-gradient RL such as GRPO).
Scope: The paper proposes and evaluates CoGR, a framework that trains two LLM keyword generators — one for queries and one for items — and co-evolves them with reinforcement learning to optimize retrieval F₁ directly.
What This Paper Is About
Retrieval systems must select a candidate set of items for a query before downstream ranking or ad auctioning happens, and errors made here cannot be undone later. Recent LLM-based approaches typically use the LLM only to rewrite or expand the query, while a separate downstream retriever still does the actual matching. CoGR asks whether LLMs can instead generate the retrieval representations for both sides — turning each query and each item into a compact set of keywords that are matched directly through an inverted index — and whether those two generators can be trained to co-adapt so that relevant query–item pairs reliably overlap.
Key Contributions
-
A two-sided generative keyword retrieval framework. CoGR trains separate query-side and item-side LLM generators whose outputs are used directly as retrieval representations, matched via keyword overlap and ranked by BM25 — preserving compatibility with existing keyword-based (inverted index) infrastructure.
-
A two-stage training pipeline. A supervised fine-tuning stage builds an aligned keyword space by deriving query-side targets from the keywords of relevant items, followed by a co-evolving reinforcement learning stage that alternately optimizes each generator with GRPO against a frozen index built by the opposite side.
-
A counterfactual marginal reward for the item side. The item generator is not rewarded on an item-centric objective; instead it receives the change in aggregate query-side F₁ caused by replacing only its own keyword set, which isolates its contribution and lets unaffected queries cancel out for efficient computation.
-
Empirical validation against 10 baselines. CoGR achieves the best overall F₁ on both an internal APP Marketplace dataset and the public WANDS benchmark, with ablations showing that co-evolving both sides, using separate generators, and SFT initialization each matter.
Main Findings
- Best F₁ on both datasets. CoGR-4B reaches F₁ of 0.3963 on the Internal APP Marketplace dataset and 0.6819 on WANDS. The paper reports improvements over the strongest baseline of 10.9% and 36.1%, respectively.
- ANCE-Qwen4B is the strongest baseline. Dense retrieval is comparatively stable across the two datasets, with ANCE-Qwen4B at F₁ 0.3575 (Internal) and 0.5012 (WANDS), where it has P 0.3756 / R 0.4354 and P 0.6136 / R 0.6092 respectively.
- Sparse and generative baselines degrade unevenly. Sparse retrieval performs substantially worse on the more challenging Internal dataset, while generative retrieval baselines degrade notably on the smaller and simpler WANDS dataset.
- Co-evolution is necessary. Variants that freeze the item-side parameters and optimize only the query-side generator (CoGR at 1.7B and 4B, and DeepRetrieval 4B) are far behind the full model: on Internal, frozen-item CoGR-4B scores F₁ 0.2617 versus 0.3963 for the full model, and on WANDS 0.4662 versus 0.6819.
- Smaller model also improves substantially. CoGR-1.7B reaches F₁ 0.3527 on Internal and 0.5685 on WANDS, both above the frozen-item versions of the same model (0.2399 and 0.3664).
- Stable co-evolution with the largest gain early. Validation F₁ rises from approximately 0.16 before co-evolving RL to approximately 0.40 after five rounds of alternation, with the largest gain in the first round and smaller but steady improvements afterward.
- Every ablated design choice hurts. On Internal with Qwen3-4B, replacing the marginal item reward with a transposed item-centric F₁ gives 0.3743, using a shared generator for both sides gives 0.3798, and removing the SFT stage gives 0.3751, all below the full model's 0.3963.
- Keywords become more specific. Between the post-SFT vocabulary and the vocabulary after five RL rounds, unigrams drop from 37% to 13% of distinct n-grams while phrases of three or more words rise from 12% to 31%. Removed keywords tend to be generic terms such as "mobile" and "fun"; added ones are longer and more informative.
- The two keyword spaces converge. The item-side vocabulary contracts during early RL rounds while the query-side vocabulary expands, and the two spaces reach a similar number of unique keywords after approximately three to four rounds.
- More input context helps. Removing item descriptions lowers F₁ from 0.3963 to 0.3759; augmenting the query-side prompt with search results from the existing search system raises F₁ to 0.4379 (P 0.4381, R 0.5002), with gains attributed to resolving ambiguous, misspelled, entity-centric, or non-English queries.
Methodology in Plain English
CoGR keeps classical keyword matching but lets LLMs decide what the keywords are.
Step 1 — Supervised initialization. The base LLM first generates M = 10 keywords per item. For each query, the paper pools the keywords of its known relevant items and keeps the N = 15 most frequent ones as the query-side target. Both generators are then fine-tuned on these targets, which guarantees the two sides start with overlapping vocabulary — enough initial recall for reinforcement learning to have a meaningful signal.
Step 2 — Co-evolving reinforcement learning. Training then alternates between the two sides, and crucially the index built by the opposite side is frozen during each update, so neither generator chases a moving target.
- Query side: the generator samples keyword sets for a query, retrieves items by keyword overlap (ranked with BM25), and receives the resulting F₁ against the labeled relevant set as reward. Sets exceeding K_max = 30 keywords get reward 0.
- Item side: the generator samples keywords for one item, and the paper builds a counterfactual index that swaps in only that item's keywords. The reward is the difference between total F₁ under the counterfactual index and total F₁ under the reference index — the item's marginal contribution to overall retrieval quality. Oversized sets get reward −1 instead of 0.
Because only one item changes per rollout, the paper caches per-query counts (retrieved items, true positives, relevant items) at the start of each item-side round, then identifies the set of queries whose retrieval status actually changes and updates only those F₁ scores. This avoids rebuilding the whole index and re-running retrieval for every rollout.
Setup. Generators use Qwen3-4B-Instruct (CoGR-4B) and Qwen3-1.7B (CoGR-1.7B). GRPO training runs on 8 NVIDIA B200 GPUs with learning rate 10⁻⁶, rollout.n = 8, max response length 512, and 5 alternating rounds (10 query-side epochs and 5 item-side epochs per round). Sampling uses temperature 1.0 / top-p 1.0 for RL rollouts and temperature 0.7 / top-p 0.8 / top-k 20 for initialization, evaluation, and indexing. The framework is noted as compatible with other RL algorithms such as PPO and with a weighted F-measure instead of F₁.
Data. The Internal APP Marketplace dataset has 13,500 / 1,500 train/evaluation queries over 39,600 applications with approximately 1,000 relevant items per query. WANDS has 430 / 50 queries over 42,994 products with approximately 200 relevant items per query. Splits are along the query dimension only, with the full item universe retained, so validation measures generalization to unseen queries. Relevance labels are binarized: Internal uses an internal LLM-as-a-judge on a five-level scale (excellent, good, acceptable, poor, bad) with "acceptable or better" treated as relevant; WANDS uses human three-level labels (Exact, Partial, Irrelevant) with Exact and Partial treated as relevant.
Why This Matters
The work pushes back on a common design assumption: that LLM reasoning should only improve the query, leaving the matching representation to a fixed retriever. By making both sides generative and then letting them co-adapt under a single retrieval objective, CoGR shows a path to better alignment without abandoning the inverted index that production keyword systems already rely on — a practical constraint especially relevant in sponsored search, where advertisers bid directly on keywords.
Real-world applications:
- Sponsored search and ad retrieval, where keyword-based matching and inverted indexes are already standard and bidders bid on keywords.
- App store and marketplace search, the internal dataset setting, where a query must retrieve many relevant apps or products.
- Product search over catalog titles and descriptions, as tested on the public WANDS/Wayfair dataset.
- Query understanding for ambiguous or misspelled inputs, where the ablation shows that adding existing search results to the query-side prompt raises F₁ from 0.3963 to 0.4379.
Industry relevance: The paper is a collaboration between the University of North Carolina at Chapel Hill and Apple, and evaluates on a confidential internal APP Marketplace dataset using Qwen backbones and the verl RL framework. The approach's compatibility with existing keyword infrastructure, its use of open-weight models at 1.7B and 4B scale, and the paper's explicit framing of precision as tied to user experience and recall as tied to revenue suggest an intended path toward deployment in commercial retrieval stacks.
Future Directions
- Rewards beyond relevance. The authors propose extending the reward from F₁ to downstream business objectives such as irrelevant ads percentage and revenue gain.
- Stronger ranking over generated keywords. The current system ranks retrieved items with BM25 over the generated keywords; the paper identifies designing better ranking as an open direction.
- Algorithmic substitutes. The reward design is stated to work with other RL algorithms such as PPO, and with a weighted F-measure when precision and recall have different business priorities — leaving those variants untested here.
- Vocabulary balance and scale. The analysis shows item- and query-side vocabularies converging only after roughly three to four rounds, and evaluation is limited to two datasets and two model scales; whether the co-evolution dynamics hold at larger scale or on datasets with sparser relevance annotation is left open.
Target Audience
Researchers and practitioners in information retrieval, search, and recommendation who are interested in LLM-driven retrieval representations, reinforcement learning with verifiable rewards, or replacing dense and generative retrieval with learned lexical representations. It is also relevant to industrial search and advertising engineers who need improvements that remain compatible with existing inverted-index infrastructure, and to graduate students studying co-training or self-evolving multi-agent LLM systems.
Authors’ abstract
Retrieval is the first stage of modern search and advertising systems, selecting a candidate set from a large item universe for downstream ranking and auction. Recent work increasingly leverages LLMs to improve retrieval through query expansion, data synthesis, and retrieval-feedback training. However, the generative component is typically used for query-side augmentation, while final matching is still delegated to a downstream retriever. We introduce CoGR, a retrieval framework that instead trains LLMs to directly construct retrieval representations on both query and item sides. Each generator produces a compact set of keywords, which are matched directly through an inverted index, preserving compatibility with existing keyword-based retrieval infrastructure. CoGR uses a two-stage training pipeline. Supervised fine-tuning first establishes an aligned keyword space, after which co-evolving reinforcement learning alternately optimizes the query- and item-side generators with GRPO against the opposite side's frozen index. Both sides optimize the same query-to-item retrieval $F_1$ objective: the query side receives retrieval $F_1$ directly, while the item side receives a counterfactual marginal reward measuring the change in query-side $F_1$ caused by its generated keywords. Across 10 representative sparse, dense, and generative baselines, CoGR achieves the best performance on both an internal APP Marketplace dataset and the public WANDS benchmark, improving $F_1$ over the strongest baseline by $10.9\%$ and $36.1\%$, respectively. Further analysis shows stable co-evolution and increasingly aligned query--item keyword spaces over training.