Nautil-RLVR

A LoRA adapter for Qwen3.5-9B that investigates a case file by requesting evidence item by item, and then either closes the case with a conclusion grounded in what it read (CASE CLOSED) or leaves it open and says what is missing (CASE NOT CLOSED). It continues Nautil-SFT with reinforcement learning on the closure decision.

📄 Paper: Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case (arXiv link coming soon) 💻 Code: etigerstudio/Nautil 📚 Dataset: etigerstudio/Nautil 🧭 Demo: etigerstudio/Nautil-Demo

Training

GRPO starting from Nautil-SFT. The reward checks only the closure decision, by program, with no judge: +1 when the decision matches the label, −1 when it does not or the closure marker is missing, plus a small format penalty. Cases are drawn uniformly from the eight source × label strata so that no source dominates the updates. 22 cases × 8 rollouts per step, learning rate 5e-6, no KL term; this checkpoint is step 30 of 50.

The step was chosen on a screening set that included 22 of the 64 test cases and 20 OOD cases; Nautil-SFT was selected before any test-set use.

Results (from the paper)

Test set: 64 non-host cases (23 should close, 41 should not); balanced accuracy and its two sides, in %.

Should close → closed Should not → not closed Balanced Within-source
Qwen3.5-9B (base) 81.2 18.7 49.9 48.2
Abstain-R1 (3B) 8.7 47.2 27.9 25.3
Gemini 3.8 Flash¹ 100.0 58.5 79.3 77.0
Nautil-SFT 44.9 93.5 69.2 60.4
Nautil-RLVR 81.2 85.4 83.3 74.1
Source-only rule 78.3 87.8 83.0 50.0

¹ One sample per case; all other models are averaged over 3 samples per case. Abstain-R1 is the closest prior work (abstention with gap explanations, single-turn RL); here it runs the same multi-turn protocol.

On the whole test set it ties the source-only rule (83.3 vs 83.0); within each source, where that rule scores 50, it reaches 74.1. It also removes missing closure markers and shortens investigations (5.5 to 4.6 turns).

The cost is evidence dependence: removing the grounds of a conclusion lowers its closure rate by 15.9 points relative to a matched control, against 26.1 for Nautil-SFT, and it addresses alternative explanations less often (17% vs 38%). Use Nautil-SFT when closures must track the evidence most closely.

How to use

The adapter is trained on the text backbone of Qwen3.5-9B (revision c202236235762e1c871ad0ccb60c8ee5ba337b9a), so load the base model as a causal LM with its text_config:

import json, torch
from transformers import AutoConfig, AutoTokenizer, Qwen3_5ForCausalLM
from peft import PeftModel

base_id = "Qwen/Qwen3.5-9B"
tokenizer = AutoTokenizer.from_pretrained(base_id)
config = AutoConfig.from_pretrained(base_id).text_config
base = Qwen3_5ForCausalLM.from_pretrained(base_id, config=config, torch_dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(base, "etigerstudio/Nautil-RLVR")

system = open("investigator_prompt.txt").read()        # files in this repository
tools = json.load(open("request_evidence_tool.json"))
messages = [{"role": "system", "content": system},
            {"role": "user", "content": case_record_and_evidence_index}]
prompt = tokenizer.apply_chat_template(messages, tools=tools, tokenize=False,
                                       add_generation_prompt=True, enable_thinking=False)

The model works in turns: it writes a short working note, calls request_evidence with up to 12 evidence IDs, receives their text as a tool message, and finally answers with CASE CLOSED or CASE NOT CLOSED. The investigation loop (serving evidence, turn limits, parsing) is in the code repository.

vLLM. vLLM loads Qwen3.5 with a language_model prefix in the weight names. vllm/ holds the same adapter with renamed weights (values unchanged); this is the copy used for every number in the paper. Generation used enable_thinking=False.

Limitations

  • Source shortcut. In this data the source of a case largely predicts whether it should close; a rule that reads only the source reaches 83.0 balanced accuracy on the test set. Closure accuracy should always be read next to that rule and within each source.
  • Host incidents. On server-incident cases no model, including frontier models, separates should-close from should-not-close cases.
  • Out of distribution. Gains on four unseen domains mostly reflect the direction of the decision; within individual domains no model is clearly above chance.
  • Gaps. When it leaves a case open, the named gap is often generic.
  • English case files only; conclusions are model outputs and must not replace an official investigation.

Citation

@article{{bi2026nautil,
  title  = {{Not Until the Evidence Says So: Teaching {{LLM}} Investigators When to Close a Case}},
  author = {{Bi, Tingzhu and Wang, Ping and Ma, Meng}},
  year   = {{2026}}
}}
Downloads last month
40
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for etigerstudio/Nautil-RLVR

Finetuned
Qwen/Qwen3.5-9B
Adapter
(759)
this model

Space using etigerstudio/Nautil-RLVR 1