Skip to content
AI.info

Research

Beyond Top-$k$ Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents

Overview Research area: Large language model (LLM) agents, specifically external skill/tool retrieval and routing over large skill registries, combined with diversity-aware subset selection using Dete

Beyond Top-$k$ Skill Retrieval: Diversity-Aware Skill Routing for LLM Agents
arXiv
2609.05824
Published
2026-09-05
Authors
Wang Wei, Tiankai Yang, Samyadeep Basu, Hongjie Chen, Yue Zhao, Zhengzhong Tu, Xiyang Hu, Franck Dernoncourt, Ryan A. Rossi, Hoda Eldardiry

AI summary

Overview

  • Research area: Large language model (LLM) agents, specifically external skill/tool retrieval and routing over large skill registries, combined with diversity-aware subset selection using Determinantal Point Processes (DPPs).
  • Technical level: Intermediate. The paper assumes familiarity with embedding-based retrieval, retrieve-and-rerank pipelines, cosine similarity, and top-k ranking; the DPP and query-residual kernel are explained from first principles.
  • Scope: The paper proposes Diverse Skill Routing (DSR), a reranking framework that selects a non-redundant shortlist of skills for a query, and evaluates it against the SkillRouter pipeline on the SkillRouter benchmark of 75 expert-verified queries over approximately 80K candidate skills.

What This Paper Is About

LLM agents increasingly load external "skills" (instructions, scripts, examples, reference documents) into context, but when a registry contains thousands or tens of thousands of skills, the router must pick only a few. Existing skill routers score each candidate independently by query relevance, which can return several near-duplicate skills and waste the limited context window while missing other skills a multi-step task needs. The paper's goal is to treat skill routing not only as relevance ranking but as complementary set selection, balancing relevance with non-redundancy.

Key Contributions

  1. Reframing skill routing as diversity-aware subset selection, motivated by redundancy in large skill registries and the compositional structure of multi-skill agent tasks.
  2. DSR (Diverse Skill Routing), a DPP-based reranking framework that combines query-dependent skill relevance (quality) with inter-skill non-redundancy in a single selection objective.
  3. A query-residual diversity kernel that removes the query-aligned component from each skill embedding before measuring inter-skill similarity, so that skills which are similar only because they are both relevant to the same query are not penalized as redundant.
  4. Empirical results on the SkillRouter benchmark showing improved recall and full coverage over a strong pointwise reranking baseline, with larger gains on multi-skill queries, plus ablations showing the query-residual kernel is the critical design choice.

Main Findings

  • Overall coverage gains at the shared cutoff: On all 75 queries, DSR raises Recall@20 from 0.754 to 0.768 and Full Coverage@20 from 0.560 to 0.573 relative to SkillRouter.
  • Larger gains for longer shortlists: At cutoff 50, DSR improves overall Recall from 0.754 to 0.808 and Full Coverage from 0.560 to 0.633. Because the released SkillRouter output is truncated at 20 ranked skills, its Recall@50 and Full Coverage@50 equal its @20 values; the paper therefore treats @20 as the shared-cutoff comparison.
  • Multi-skill queries benefit most: On the 51 multi-skill queries, DSR improves Recall@20 from 0.704 to 0.739 and Full Coverage@20 from 0.458 to 0.492. At cutoff 50, Recall goes from 0.704 to 0.773 and Full Coverage from 0.458 to 0.551.
  • The query-residual kernel is the decisive component: With reranker quality, replacing the standard cosine kernel with the query-residual kernel raises multi-skill Full Coverage@10 from 0.254 to 0.441. Naive diversity based on raw inter-skill cosine similarity reaches only 0.254 (reranker quality) and 0.178 (embedding quality), both below the pointwise baseline's 0.424.
  • Quality scores still matter: With the query-residual kernel, swapping embedding quality for reranker quality lifts Recall@10 from 0.618 to 0.711 and multi-skill Full Coverage@10 from 0.237 to 0.441. Diversity-aware selection depends on a reliable relevance signal.
  • Minimal loss in early precision: DSR's MRR@10 is 0.784 versus 0.788 for SkillRouter on all queries, and 0.788 versus 0.792 on multi-skill queries, so the coverage gains do not come from a large drop in first-relevant-skill position.
  • Single-skill queries show no difference: Top-k retrieval, SkillRouter, and DSR all achieve Recall@10 and Full Coverage@10 of 0.875 on single-skill queries, indicating pointwise relevance signals are sufficient when only one skill is required.
  • Diversity-aware selection beats a zero-shot LLM reranker: On multi-skill queries, DSR reaches Recall@10 of 0.668 and Full Coverage@10 of 0.432, compared with 0.616 and 0.331 for a zero-shot Qwen3-8B ranker, and 0.659 and 0.424 for SkillRouter. Top-k retrieval without reranking gets 0.630 and 0.381.

Methodology in Plain English

The approach starts from a standard two-stage pipeline. First, a lightweight encoder retriever embeds the query and every skill, scores pairs by cosine similarity of L2-normalized embeddings, and pulls out the top 50 candidates from the full registry. Second, a learned pointwise reranker (SR-Rank-0.6B) reads the full query-skill pair and outputs a relevance logit, which is squashed through a sigmoid to give each candidate a non-negative quality score. DSR changes only this final step, leaving the retriever and reranker untouched, so any improvement is attributable to selection rather than to a stronger model.

Instead of simply returning the highest-scoring candidates, DSR builds a DPP kernel over the retrieved candidate set, where each entry multiplies the two skills' quality scores by a query-conditioned similarity term. The determinant of a submatrix of this kernel is large when the chosen skills are both high-quality and mutually different, so maximizing that determinant picks a good, non-redundant shortlist.

The distinctive part is how similarity is measured. A plain cosine kernel treats any closeness between two skills as redundancy, which is wrong when two skills are close simply because both answer the same request. DSR instead subtracts the query-aligned component from each skill embedding to get a residual vector, then mixes that residual with the original embedding using a coefficient lambda of 0.85 and renormalizes. Cosine similarity in this residual space, rescaled to the range 0 to 1, becomes the kernel's similarity term. This penalizes overlap in what skills do, while softening penalties that come only from shared relevance to the query.

Because exactly maximizing the determinant is expensive, DSR uses greedy MAP selection with incremental Cholesky updates: it repeatedly adds the candidate with the largest marginal gain in log-determinant, stopping at k skills. The first greedy step provably selects the highest-quality skill, matching the pointwise reranker's top pick. The final selected skills are then re-sorted by quality score to produce the ranked output. Experiments ran on NVIDIA A100 80GB GPUs, evaluating cutoffs k of 10, 20, and 50. Metrics are Recall@k (fraction of target skills recovered) and Full Coverage@k (whether all target skills appear in the top k), the latter being stricter and more important for multi-skill workflows.

Why This Matters

Impact on research: The paper separates the selection objective from the retrieval and scoring stages, showing that the selection layer remains important even when the retriever and reranker are held fixed. It also demonstrates that diversity methods cannot be lifted unchanged into query-conditioned settings: the same quality scores combined with a raw cosine kernel perform worse than the pointwise baseline (multi-skill Full Coverage@10 of 0.254 or 0.178 versus 0.424). This gives a concrete negative result and a concrete fix, and connects skill routing to a long line of DPP work in summarization, recommendation, and information retrieval.

Real-world applications:

  • Enterprise agent platforms that expose thousands of internal procedural skills and must fit a small, complementary set into a limited context window.
  • Multi-step data workflows that combine spreadsheet parsing, column cleaning, data transformation, and visualization skills for a single request.
  • Cost- and latency-sensitive agent deployments where loading redundant skills wastes tokens and distracts execution.
  • Agent frameworks built on large public skill registries, where many skills overlap functionally and distractor skills are present (the benchmark's Hard tier includes 79,141 candidates with 780 LLM-generated distractors).

Industry relevance: The routing layer is a practical cost lever: DSR does not require a larger retriever or reranker, only a different final selection step over the top 50 retrieved candidates, which keeps the overhead manageable. For organizations maintaining skill/tool catalogs at scale, it offers a way to improve coverage of required capabilities without retraining the underlying models. The paper's affiliations (Virginia Tech, USC, Adobe Research, Dolby Labs, Texas A&M, Arizona State University) reflect both academic and industrial interest.

Future Directions

  • Broader evaluation. The benchmark contains only 75 expert-verified queries; the authors call for testing diversity-aware routing on more domains, more complex workflows, and different skill types.
  • Joint retrieval and selection. DSR operates after candidate retrieval, so any required skill missing from the candidate set is unrecoverable. Improving candidate generation and diversity-aware selection together is flagged as an important direction.
  • Connect coverage to task success. The current metrics are retrieval-based (Recall, Full Coverage) and do not measure whether the agent actually completes the task; agents may still fail from incorrect tool use, poor planning, or conflicting skill instructions. The paper reports no end-to-end execution results.
  • Reduce selection overhead. The DPP step adds cost after reranking. The paper notes that the overhead may matter for latency-sensitive deployments or much larger candidate sets, and points to more efficient approximations and adaptive candidate sizes. No latency or wall-clock measurements are reported.

Target Audience

Researchers and engineers working on LLM agents, tool and skill retrieval, and retrieval-augmented systems, especially those dealing with large registries of overlapping capabilities. It is also relevant to readers interested in diversity-aware ranking and DPP-based subset selection who want to see how those methods behave under query-conditioned relevance, and to practitioners who need to fit complementary capabilities into a bounded context window without retraining their retrieval or reranking models. Readers without background in embedding retrieval and subset selection will need to consult the cited DPP literature to follow the method section fully.

Authors’ abstract

Large language model (LLM) agents increasingly rely on external skills, but routing user requests over large skill registries is difficult because many skills are functionally redundant while complex tasks often require complementary skill sets. Existing skill routers typically rank candidates independently by query relevance, which can waste context budget on redundant skills. We propose Diverse Skill Routing (DSR), a diversity-aware reranking framework that uses a Determinantal Point Process to balance relevance and non-redundancy. DSR introduces a query-residual diversity kernel that penalizes redundant skill overlap while reducing penalties caused only by shared query relevance. On the SkillRouter benchmark, DSR improves recall and full coverage over a strong pointwise reranking baseline, with larger gains on multi-skill queries. These results suggest that skill routing should be treated not only as relevance ranking, but also as complementary set selection.

Read the original paper