Research
PatchHolmes: Agentic Patch Retrieval via Listwise Selection
Overview Research area: Information retrieval (cs.IR) applied to software security — specifically patch retrieval, the task of mapping a CVE to the commit that fixes it in an upstream repository. Tech

- arXiv
- 2609.38807
- Published
- 2026-09-30
- Authors
- Guanqun Yang, Yingming Zhou, Jiangrui Zheng, Shudong Hao, Xueqing Liu
AI summary
Overview
- Research area: Information retrieval (cs.IR) applied to software security — specifically patch retrieval, the task of mapping a CVE to the commit that fixes it in an upstream repository.
- Technical level: Advanced. The paper assumes familiarity with retrieval metrics (Recall@K, NDCG, MRR), rank fusion, LLM agents and tool-use loops, though the core idea is explainable without that background.
- Scope: The paper introduces PatchHolmes, a two-phase patch retrieval system that combines a hybrid first-stage retriever with an LLM agent that reads a 100-candidate list "listwise" through four budgeted tools, and evaluates it against Favia, IRCoT, SPFinder and PatchFinder on two CVE benchmarks.
What This Paper Is About
Every disclosed software vulnerability needs to be linked to the commit that fixes it, but a large share of CVE records have no such patch link: the paper cites figures of 60% to 63% of CVEs in the major advisory databases missing their patch link, and notes that auxiliary-information matching recovers only 12% to 53% of disclosed open-source vulnerabilities even with manual assistance. Patch retrieval systems exist, but the paper argues they are limited in two ways: they are open-loop (score every candidate once, emit a ranking, never look closer) and they rely on pre-LLM encoders whose 512-token window cannot fit diffs that average about 15,000 tokens in this corpus. PatchHolmes' goal is to close both gaps by letting an LLM agent survey the full shortlist and selectively read individual commits before submitting a single best answer.
Key Contributions
-
A two-phase listwise system. PatchHolmes pairs a hybrid Phase 1 retriever — BM25 with time decay fused with a Qwen3-Embedding-8B dense retriever via Reciprocal Rank Fusion at k=60 — into a top-100 candidate set, with a Phase 2 agent that inspects that set listwise rather than scoring each candidate independently.
-
A four-tool agentic inspection loop. The Phase 2 agent drives
list_candidates,read_commit,read_file_diff, andsubmit_answerover up to 15 iterations, typically reading 3 to 10 of the 100 candidates before submitting one commit per CVE. The paper argues this is the minimal sufficient tool set for the task: a survey call, two granularities of commit reading (8,000 characters at commit level, 16,000 for a single-file drill-down), and a submit. -
Measured head-to-head gains over pointwise and retrieve-and-CoT baselines. On GitHubAD, PatchHolmes beats the pointwise binary classifier Favia by 25.34% Recall@1 and IRCoT by 31.40%, at one agent conversation per CVE versus Favia's ten.
-
Demonstrated transfer and backbone-independence. The same agent transferred unchanged to the PatchFinder_top10 pool lifts Recall@1 from PatchFinder's own top-1 pick (24.28%) to 39.86%, and swapping the LLM backbone within the Qwen family changes Recall@1 by under 1%, with the whole system running on a frozen open-weight model over a local Git repository, without fine-tuning or external search APIs. The system is open-sourced at https://github.com/Aizhouym/PatchHolmes.
Main Findings
-
PatchHolmes leads every metric on GitHubAD. With the Qwen3-235B backbone over the 809-CVE working subset, Recall@1 is 59.95% for PatchHolmes versus 34.61% for Favia, 28.55% for IRCoT, and 27.32% for SPFinder. At Recall@10 the margins narrow: PatchHolmes 77.26%, Favia 72.31%, SPFinder 70.36%.
-
The agent, not the retriever, explains the gap. Giving every method PatchHolmes' own Phase 1 candidates, taking Phase 1's rank-1 candidate with no agent reaches 32.63% Recall@1, while the agent on the same 100 candidates reaches 59.95% — a gain of 27.32% from the agent alone. With candidates matched, Favia rises from 34.61% to 39.80% and IRCoT from 28.55% to 37.58%, leaving remaining gaps of 20.15% and 22.37% attributable to the selection method.
-
Pointwise scoring loses precision at the top. Favia's Recall@10 is only 4.95% behind PatchHolmes (72.31% vs 77.26%), but its Recall@1 is 25.34% behind. SPFinder shows the same shape. The paper attributes this to independent per-candidate scoring leaving no joint signal to break ties when several candidates look plausible, illustrated by CVE-2015-9251, where Favia gives all 10 candidates identical confidence 5 and picks the rank-1 refactor rather than the gold commit at rank 9.
-
The benefit transfers to a pool the authors did not build. On PatchFinder_top10 (1,252 CVEs, 10 pre-filtered candidates each, Phase 1 skipped), the agent lifts Recall@1 from 24.28% to 39.86% — an absolute gain of 15.58% and a relative gain of 64.16%. Because only 577 of the 1,252 CVEs have the gold fix anywhere in the supplied pool, Recall@10 caps at 46.10%; PatchHolmes recovers 71.40% of the gap to that ceiling.
-
The dense retrieval path carries most of Phase 1, but fusion still helps. Ablating Phase 1 paths with the unchanged agent: BM25 only reaches 39.68% Recall@1, BM25 with time decay 49.57%, dense-only 57.11% (within 2.84% of the full system). RRF on top adds +2.84% Recall@1 and +5.32% Recall@10.
-
Listwise viewing is the single largest tool contribution. On fixed Phase 1 candidates, the ladder from no agent (32.63% Recall@1) to
list_candidatesplussubmit_answer(48.21%, +15.58%) to file-level overview without diff text (57.48%, +9.27%) to full diff reading (59.95%, +2.47%) shows each tool adds recall. -
Richer CVE metadata does not help. Replacing the short CVE description with the full NVD report (mean 1,786 vs about 150 characters) leaves Recall@1 flat (59.83% vs 59.95%), drops Recall@10 by 2.11%, and raises the context-overflow error rate from 0.49% to 2.60%.
-
Backbone scale matters less than the loop. Within the Qwen family, Recall@1 is 59.95% on Qwen3-235B and 59.09% on Qwen3-Coder-30B, a 0.86% gap. A second family stays above the no-agent floor: gpt-oss-120B at 56.98% and gpt-oss-20B at 49.81%, with the latter held down partly by lower submission rates at the 15-iteration cap (70% and 54%, versus 90.5% for Qwen3-235B).
-
Cost favors PatchHolmes. With Qwen3-235B the agent consumes a mean of 96,118 input tokens and 846 output tokens per CVE in about 8 tool calls, inspecting 5 commits in 86 seconds (median 49); about USD 0.007 per CVE and USD 58 for a pass over the full 8,401-CVE corpus, against about USD 400 for Favia at the same rate. All gaps are reported as statistically significant, with McNemar's exact test giving p = 2.4 × 10^-44 for PatchHolmes versus Favia at Recall@1.
-
An external sanity check matches the benchmark. Running the full system on 35 independent real-world CVEs from 11 repositories disclosed between 2013 and 2025 reached Recall@1 of 60.0%, close to the 59.95% on GitHubAD.
Methodology in Plain English
The system works in two stages that share nothing except a list of candidate commit identifiers.
Stage one — cast a wide net. For a given CVE description and target repository, two retrieval paths run in parallel. One is a lexical path: BM25 scoring over the commit message and the diff body (weighted 0.35 and 0.15, each min-max normalized against the top 10,000 hits), plus two time-rank signals (weighted 0.30 and 0.20) that favor commits whose timestamps sit close to the CVE's reserve date and public-disclosure date. The time score takes the form 1/(1+2|r−c|), peaking at 1 for the closest commit and decaying symmetrically. The other path embeds the CVE description and every commit into a shared 4096-dimensional space with Qwen3-Embedding-8B and ranks by cosine similarity; diffs are packed into a 6,000-character budget by sorting lines into four information-density buckets (file paths, hunk headers, ± lines, context lines) and filling from the highest priority down, keeping whole lines only. The two rankings are fused with Reciprocal Rank Fusion at k=60, and the top 100 are handed to stage two.
Stage two — read selectively and commit to one answer. An LLM agent, instantiated in the OpenHands-SDK Conversation runtime, receives the CVE description verbatim and works through four tools. list_candidates returns one line per candidate in Phase 1 rank order (rank, commit id, first line of the message, file-change tally) and is the only call that shows all 100 at once. read_commit returns the commit message, a file manifest, and the diff under an 8,000-character budget; the diff is parsed into per-file records, each file is assigned a priority that rewards source-code tags over test or doc tags, larger edits up to a cap, and path-token overlap with the CVE description, while penalizing refactor-shaped diffs, and files are then rendered in descending priority, compressed if needed, or pointed to via read_file_diff. read_file_diff fetches a single file's diff under a 16,000-character budget for cases where the commit-level view truncated something relevant. submit_answer(commit_id, reasoning) ends the conversation exactly once, and the submitted identifier is the retrieved patch. The agent runs with a maximum of 15 iterations and stuck detection, terminating early if it submits, hits the budget, or repeats an action twice.
Evaluation. The authors compare against Favia (pointwise LLM binary classifier with two metadata tools), IRCoT (the canonical retrieve-then-reason baseline run through FlashRAG), and SPFinder (a non-agentic two-phase retriever with a hierarchical embedding and gradient-boosted reranker, no LLM calls at inference). All LLM-based methods share the Qwen3-235B backbone in the head-to-head so the model is not a confounder. Metrics are Recall@K, NDCG@K and MRR per CVE, with runtime errors counted as misses rather than dropped.
Why This Matters
Impact on research. The paper's central evidence is that on fixed candidate sets the agent loop adds 27.32% Recall@1 while a roughly 7x change in backbone active parameters moves it by under 1% within a family. It also notes the finding is consistent with prior work reporting that swapping the retriever changes end-to-end accuracy more than swapping the agent does. That reframes patch retrieval — and arguably other single-answer retrieval tasks — as a selection-interface problem rather than a model-scale problem, and it supplies a working, open-source artifact for measuring that.
Real-world applications.
- Security advisories and CVE records: automatically populating the patch links that 60% to 63% of CVEs currently lack.
- CVSS scoring and affected-version-range tracking: knowing the exact fix commit lets tools compute which released versions are patched.
- SBOM-driven supply-chain scanners: determining whether a shipped dependency already contains the fix.
- Maintainer backlogs: systems such as the NVD face curation delays, and the paper cites CVE-2013-1814 in Apache Rave, whose patch was not surfaced until a decade after disclosure.
Industry relevance. The system is designed for the constraints a production deployment would face: a frozen open-weight model, no per-repository fine-tuning, no paid external search API, and offline-precomputed embeddings per repository cached as a single matrix so per-CVE retrieval is one matrix multiplication. The cost profile is the selling point — about 96,118 input tokens and about USD 0.007 per CVE, roughly 7x fewer tokens than Favia at 25.34% higher Recall@1, and about USD 58 versus about USD 400 for a full 8,401-CVE pass.
Future Directions
- Calibrated confidence scores. The system's output is described as a ranked list, not a single guess, with the inspected candidates placed above the remaining Phase 1 order. Producing calibrated confidence scores for that list is explicitly stated as future work.
- The patch-link sparsity blind spot. Because only CVEs with an already-known fix can be scored, every benchmark in this area — including this one — evaluates on the 37% to 40% of the population whose patch has been located. The hard cases where no one has ever found the fix are absent from all of them, and the paper states this scope limit is not removed by its own 35-CVE external check.
- Multi-fix ambiguity. 7.28% of the 8,401 source CVEs (612) have more than one recorded fix commit, and backports, split refactors and downstream merges may each be legitimate fixes that were never recorded. A different-but-valid pick currently scores as a miss for PatchHolmes and every other single-pick method.
- Benchmark ceilings and coverage. On PatchFinder_top10 the gold fix is present for only 577 of 1,252 CVEs, capping any single-pick selector at Recall@1 = 46.10%; improving either the candidate pool or the dataset itself is left open. The authors also note their backbone comparison covers open-weight models from 20B to 235B total parameters in two families and should not be read as proof that backbone scale is irrelevant.
Target Audience
This paper is most useful to researchers and engineers working on retrieval-augmented LLM agents, vulnerability management automation, and code search — particularly those who need to select one answer from a fixed candidate list rather than rank a large open corpus. It is also relevant to practitioners building CVE-to-commit mapping pipelines who care about token cost per query and about running on frozen open-weight models with no external search API. Readers without a retrieval or security background will find the framing accessible but the evaluation dense.
Authors’ abstract
Patch retrieval, the task of finding the commit that fixes a known vulnerability, is the foundation of vulnerability management workflows, yet 60% to 63% of CVEs in the major advisory databases lack a patch link. We present PatchHolmes, a two-phase patch retrieval system that pairs a hybrid first-stage retriever with an agentic second-stage inspection loop. Unlike pointwise prior work that scores each candidate independently, the Phase 2 agent reads the top-100 listwise: it sees the full candidate list at once and selectively reads 3 to 10 commits through four budgeted tools before submitting a single best commit. On GitHubAD, PatchHolmes beats the pointwise binary classifier Favia by 25.34% Recall@1 and the retrieve-and-CoT baseline IRCoT by 31.40%, at one agent conversation per CVE versus Favia's ten; with the candidate set held identical, the agent adds 27.32% Recall@1 over taking the retriever's top candidate, and the same agent, transferred unchanged to PatchFinder_top10, lifts Recall@1 from PatchFinder's own top-1 pick (24.28%) to 39.86%. Swapping the LLM backbone within the Qwen family changes Recall@1 by under 1%, and a second model family (gpt-oss) stays far above the no-agent floor, so the gain comes from the listwise agent loop; the entire system runs on a frozen open-weight model over a local Git repository, without fine-tuning or external search APIs.