Instructions to use etigerstudio/Nautil-RLVR with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use etigerstudio/Nautil-RLVR with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-9B") model = PeftModel.from_pretrained(base_model, "etigerstudio/Nautil-RLVR") - Notebooks
- Google Colab
- Kaggle
Nautil-RLVR
A LoRA adapter for Qwen3.5-9B that investigates a case file by requesting evidence item by item, and then either closes the case with a conclusion grounded in what it read (CASE CLOSED) or leaves it open and says what is missing (CASE NOT CLOSED). It continues Nautil-SFT with reinforcement learning on the closure decision.
📄 Paper: Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case (arXiv link coming soon) 💻 Code: etigerstudio/Nautil 📚 Dataset: etigerstudio/Nautil 🧠Demo: etigerstudio/Nautil-Demo
Training
GRPO starting from Nautil-SFT. The reward checks only the closure decision, by program, with no judge: +1 when the decision matches the label, −1 when it does not or the closure marker is missing, plus a small format penalty. Cases are drawn uniformly from the eight source × label strata so that no source dominates the updates. 22 cases × 8 rollouts per step, learning rate 5e-6, no KL term; this checkpoint is step 30 of 50.
The step was chosen on a screening set that included 22 of the 64 test cases and 20 OOD cases; Nautil-SFT was selected before any test-set use.
Results (from the paper)
Test set: 64 non-host cases (23 should close, 41 should not); balanced accuracy and its two sides, in %.
| Should close → closed | Should not → not closed | Balanced | Within-source | |
|---|---|---|---|---|
| Qwen3.5-9B (base) | 81.2 | 18.7 | 49.9 | 48.2 |
| Abstain-R1 (3B) | 8.7 | 47.2 | 27.9 | 25.3 |
| Gemini 3.8 Flash¹ | 100.0 | 58.5 | 79.3 | 77.0 |
| Nautil-SFT | 44.9 | 93.5 | 69.2 | 60.4 |
| Nautil-RLVR | 81.2 | 85.4 | 83.3 | 74.1 |
| Source-only rule | 78.3 | 87.8 | 83.0 | 50.0 |
¹ One sample per case; all other models are averaged over 3 samples per case. Abstain-R1 is the closest prior work (abstention with gap explanations, single-turn RL); here it runs the same multi-turn protocol.
On the whole test set it ties the source-only rule (83.3 vs 83.0); within each source, where that rule scores 50, it reaches 74.1. It also removes missing closure markers and shortens investigations (5.5 to 4.6 turns).
The cost is evidence dependence: removing the grounds of a conclusion lowers its closure rate by 15.9 points relative to a matched control, against 26.1 for Nautil-SFT, and it addresses alternative explanations less often (17% vs 38%). Use Nautil-SFT when closures must track the evidence most closely.
How to use
The adapter is trained on the text backbone of Qwen3.5-9B (revision c202236235762e1c871ad0ccb60c8ee5ba337b9a), so load the base model as a causal LM with its text_config:
import json, torch
from transformers import AutoConfig, AutoTokenizer, Qwen3_5ForCausalLM
from peft import PeftModel
base_id = "Qwen/Qwen3.5-9B"
tokenizer = AutoTokenizer.from_pretrained(base_id)
config = AutoConfig.from_pretrained(base_id).text_config
base = Qwen3_5ForCausalLM.from_pretrained(base_id, config=config, torch_dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(base, "etigerstudio/Nautil-RLVR")
system = open("investigator_prompt.txt").read() # files in this repository
tools = json.load(open("request_evidence_tool.json"))
messages = [{"role": "system", "content": system},
{"role": "user", "content": case_record_and_evidence_index}]
prompt = tokenizer.apply_chat_template(messages, tools=tools, tokenize=False,
add_generation_prompt=True, enable_thinking=False)
The model works in turns: it writes a short working note, calls request_evidence with up to 12 evidence IDs, receives their text as a tool message, and finally answers with CASE CLOSED or CASE NOT CLOSED. The investigation loop (serving evidence, turn limits, parsing) is in the code repository.
vLLM. vLLM loads Qwen3.5 with a language_model prefix in the weight names. vllm/ holds the same adapter with renamed weights (values unchanged); this is the copy used for every number in the paper. Generation used enable_thinking=False.
Limitations
- Source shortcut. In this data the source of a case largely predicts whether it should close; a rule that reads only the source reaches 83.0 balanced accuracy on the test set. Closure accuracy should always be read next to that rule and within each source.
- Host incidents. On server-incident cases no model, including frontier models, separates should-close from should-not-close cases.
- Out of distribution. Gains on four unseen domains mostly reflect the direction of the decision; within individual domains no model is clearly above chance.
- Gaps. When it leaves a case open, the named gap is often generic.
- English case files only; conclusions are model outputs and must not replace an official investigation.
Citation
@article{{bi2026nautil,
title = {{Not Until the Evidence Says So: Teaching {{LLM}} Investigators When to Close a Case}},
author = {{Bi, Tingzhu and Wang, Ping and Ma, Meng}},
year = {{2026}}
}}
- Downloads last month
- 40