Papers
arxiv:2610.05107

SearchJev: A Fast and Calibrated System-1 Model for Search Agents

Published on Oct 4
· Submitted by
youganglyu
on Oct 6
Authors:
,
,
,
,
,
,
,
,

Abstract

Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates search decisions from System-2 reasoning and generation. Given a search state and a decision schema, SearchJev directly scores legal options without autoregressive output generation. We propose Soft-Label Learning for Calibrated Decisions (SLCD) to learn decision probabilities from uncertain supervision and calibrate their confidence. In a dual-system search agent, SearchJev handles short decisions and delegates uncertain judgments to System 2, which retains planning, query generation, and answer composition. We also introduce SearchDecision-Bench, a benchmark unifying six types of search decisions for training and evaluation. On SearchDecision-Bench, SEARCHJEV improves decision quality over same-size Qwen3.5 autoregressive models, achieves 5.2-5.3 times faster decisions, and reduces average expected calibration error by 41-74%. On BrowseComp-Plus, the dual-system agents achieve a 3.7-4.7 times speedup in active search time while improving answer accuracy from 45% to up to 54%.

Community

Paper submitter

SearchJev, a fast and calibrated System-1 model for search agents.

Search agents repeatedly make small but important decisions: Is this passage relevant? Is the evidence sufficient? Which link should I open next? Asking a large LLM to generate an answer for every decision adds latency, and its reported confidence can be unreliable.

Inspired by fast and slow thinking, we separate these decisions from the heavier reasoning work. SearchJev acts as System 1: it directly scores the available options and returns calibrated probabilities. High-confidence decisions stay with SearchJev; uncertain ones go to System 2, which also handles planning, query generation, and the final answer.

We evaluate both the individual decisions and their impact on the full search process:

  • SearchDecision-Bench: 5.2–5.3× faster decisions, 41–74% lower average calibration error, and better quality across all six decision types than same-size Qwen3.5 (JSON).
  • BrowseComp-Plus: 3.7–4.7× faster active search, with accuracy rising from 45% to up to 54% over System 2 alone.
  • Compared with Jev 1.13: SearchJev-4B raises BrowseComp-Plus accuracy from 46% to 54%; SearchJev-0.8B matches 46% with 15% less active search time.

🤗 Models: https://huggingface.co/SearchJev/models
💻 Code: https://github.com/EvoScientist/SearchJev

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.05107
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.05107 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.05107 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.05107 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.