AgSpec: Pushing the Limits of Retrieval-Based Speculative Decoding in Coding Agent Pipelines
Abstract
Retrieval-based speculative decoding (SD) drafts tokens by copying continuations from existing text, which suits coding agents that repeatedly reproduce code, logs, and earlier attempts. Yet existing methods fall short in agent pipelines: much of the reusable text is missing from their corpora or stored in a form that differs from what the agent emits, and their draft lengths ignore that accept length varies across agents and drifts over turns. We present AgSpec, a framework that supplies the corpus and draft-length policies that existing retrieval engines lack in coding-agent pipelines. AgSpec retrieves from session, workspace, and global corpora, retaining the ongoing session trajectory and indexing opened files in the agent's emission format. It bounds each agent's draft length with an offline-profiled cap and adapts the length online from verification feedback. On two repository-level multi-agent coding benchmarks, AgSpec outperforms five retrieval-based drafters and EAGLE-3 in most evaluated settings, raising generation throughput over autoregressive decoding up to 4.37times at batch size 1 and 4.76times at batch size 16. AgSpec also remains effective on benchmarks without a repository or a multi-agent pipeline, showing that its gains generalize to coding agents broadly.
Community
We proposed a retrieval-based speculative decoding framework, AgSpec, for coding agent pipelines. AgSpec provides the proper policy for a retrieval engine, such as corpus design and draft-length decisions. We achieve about 4 times higher throughput than AR decoding, beating retrieval-based speculative decoding methods and Eagle 3 on two repository-level coding benchmarks.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- AgentSpec: Speculative Decoding for Batch Inference of LLM Agents (2026)
- ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference (2026)
- AsymSpec: Context-Asymmetric Speculative Decoding for Agentic LLMs (2026)
- EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents? (2026)
- TokenCast: Forecasting Token Consumption During LLM Agent Execution (2026)
- AgentReplay: Token-Wise Trace Replay Is Essential for Fair Serving System Performance Benchmarking (2026)
- LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.01108 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper