Papers
arxiv:2609.38807

PatchHolmes: Agentic Patch Retrieval via Listwise Selection

Published on Sep 30
ยท Submitted by
David Yang
on Oct 1
Authors:
,
,
,
,

Abstract

Patch retrieval, the task of finding the commit that fixes a known vulnerability, is the foundation of vulnerability management workflows, yet 60% to 63% of CVEs in the major advisory databases lack a patch link. We present PatchHolmes, a two-phase patch retrieval system that pairs a hybrid first-stage retriever with an agentic second-stage inspection loop. Unlike pointwise prior work that scores each candidate independently, the Phase 2 agent reads the top-100 listwise: it sees the full candidate list at once and selectively reads 3 to 10 commits through four budgeted tools before submitting a single best commit. On GitHubAD, PatchHolmes beats the pointwise binary classifier Favia by 25.34% Recall@1 and the retrieve-and-CoT baseline IRCoT by 31.40%, at one agent conversation per CVE versus Favia's ten; with the candidate set held identical, the agent adds 27.32% Recall@1 over taking the retriever's top candidate, and the same agent, transferred unchanged to PatchFinder_top10, lifts Recall@1 from PatchFinder's own top-1 pick (24.28%) to 39.86%. Swapping the LLM backbone within the Qwen family changes Recall@1 by under 1%, and a second model family (gpt-oss) stays far above the no-agent floor, so the gain comes from the listwise agent loop; the entire system runs on a frozen open-weight model over a local Git repository, without fine-tuning or external search APIs.

Community

Paper submitter

Accepted at AACL-IJCNLP 2026, Main Conference.

Every software vulnerability needs to be paired with the commit that fixed it. Security advisories need that pairing. So do severity scores, affected-version trackers, and supply-chain scanners.

โŒ In the GitHub Advisory Database and the National Vulnerability Database, 60% to 63% of entries do not have it.

The search is hard for three reasons:

โŒ The haystack is large. A repository can carry up to 1.4 million commits, and 49% of these vulnerabilities are filed against repositories with more than 5,000.
โŒ The candidates are long. The average commit diff in our corpus runs about 15,000 tokens. The text encoders that earlier methods use read the first 512.
โŒ The words do not match. The vulnerability report names the symptom, "buffer overflow". The commit that fixes it names the bug, "out-of-bounds read".

๐Ÿ” The failure that started this paper.

The previous best agentic system scores each of 10 candidates in its own separate call: yes or no, plus a confidence.

On CVE-2015-9251, a cross-site scripting bug in jQuery, it marks all ten "yes" at confidence 5. Nothing distinguishes them, so the rank-1 pick is whatever the list happened to put first. That was an unrelated refactor of the same code path.

โœ… Our fix: let the agent see the whole list.

PatchHolmes reads the top 100 candidates at once, then opens individual commits through four budgeted tools and submits a single answer. The search narrows in four steps: up to 1.4 million commits in the repository, 100 candidates after the first stage, 5 commits the agent actually opens, 1 answer submitted. On that jQuery case it reads five commits, in list order 1, 2, 9, 4, 3, and picks the one that blocks automatic execution of JavaScript responses.

๐Ÿ“Š Results. Recall@1 is the share of vulnerabilities where the single submitted commit is the correct one, so higher is better.

โœ… PatchHolmes 59.95%, against 34.61% for the pointwise system and 28.55% for a general retrieve-and-reason baseline
โœ… On identical candidates, the agent lifts Recall@1 from 32.63% to 59.95%
โœ… The same agent, moved unchanged to another benchmark's candidate pool, lifts Recall@1 from 24.28% to 39.86%
โœ… Swapping the language model inside one family moves Recall@1 by under 1 point
โœ… About 96,000 input tokens, 8 tool calls, and 5 commits read per vulnerability

The whole system runs on a frozen open-weight model over a local Git clone. No fine-tuning and no paid search API, because both cost more as the backlog grows.

๐Ÿ™ Joint work with Yingming Zhou, Jiangrui Zheng, Shudong Hao, and Xueqing Liu at Stevens Institute of Technology.

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.38807 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.38807 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.38807 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.