Nautil-SFT

A LoRA adapter for Qwen3.5-9B that investigates a case file by requesting evidence item by item, and then either closes the case with a conclusion grounded in what it read (CASE CLOSED) or leaves it open and says what is missing (CASE NOT CLOSED).

📄 Paper: Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case (arXiv link coming soon) 💻 Code: etigerstudio/Nautil 📚 Dataset: etigerstudio/Nautil 🧭 Demo: etigerstudio/Nautil-Demo 🔁 Reinforcement-learned version: etigerstudio/Nautil-RLVR

Training

Supervised fine-tuning on 545 teacher trajectories (GPT-6 Sol) over audited cases from aviation, rail, maritime, chemical-safety and vehicle-defect reports and production server incidents. Each trajectory keeps a public working note with a hypothesis ledger before every evidence request.

  • LoRA r=16, alpha=32, dropout 0, all linear layers of the text backbone
  • 2 epochs, 546 steps in total (this checkpoint is the last step of epoch 2); 32k context, bf16
  • Loss on assistant turns only

Results (from the paper)

Test set: 64 non-host cases (23 should close, 41 should not); balanced accuracy and its two sides, in %.

Should close → closed Should not → not closed Balanced Within-source
Qwen3.5-9B (base) 81.2 18.7 49.9 48.2
Abstain-R1 (3B) 8.7 47.2 27.9 25.3
Gemini 3.8 Flash¹ 100.0 58.5 79.3 77.0
Nautil-SFT 44.9 93.5 69.2 60.4
Nautil-RLVR 81.2 85.4 83.3 74.1
Source-only rule 78.3 87.8 83.0 50.0

¹ One sample per case; all other models are averaged over 3 samples per case. Abstain-R1 is the closest prior work (abstention with gap explanations, single-turn RL); here it runs the same multi-turn protocol.

Its closures follow the evidence: when all grounds of a conclusion are removed, its closure rate drops 26 points relative to a matched control (base model: 6). Overstated conclusions fall from 97% (base) to 35%, and answers that are both right about the cause and not overstated rise from 3% to 43%.

It is cautious: it leaves most undetermined cases open, but also closes fewer than half of the cases that should close. Nautil-RLVR balances the two sides at some cost in evidence dependence.

How to use

The adapter is trained on the text backbone of Qwen3.5-9B (revision c202236235762e1c871ad0ccb60c8ee5ba337b9a), so load the base model as a causal LM with its text_config:

import json, torch
from transformers import AutoConfig, AutoTokenizer, Qwen3_5ForCausalLM
from peft import PeftModel

base_id = "Qwen/Qwen3.5-9B"
tokenizer = AutoTokenizer.from_pretrained(base_id)
config = AutoConfig.from_pretrained(base_id).text_config
base = Qwen3_5ForCausalLM.from_pretrained(base_id, config=config, torch_dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(base, "etigerstudio/Nautil-SFT")

system = open("investigator_prompt.txt").read()        # files in this repository
tools = json.load(open("request_evidence_tool.json"))
messages = [{"role": "system", "content": system},
            {"role": "user", "content": case_record_and_evidence_index}]
prompt = tokenizer.apply_chat_template(messages, tools=tools, tokenize=False,
                                       add_generation_prompt=True, enable_thinking=False)

The model works in turns: it writes a short working note, calls request_evidence with up to 12 evidence IDs, receives their text as a tool message, and finally answers with CASE CLOSED or CASE NOT CLOSED. The investigation loop (serving evidence, turn limits, parsing) is in the code repository.

vLLM. vLLM loads Qwen3.5 with a language_model prefix in the weight names. vllm/ holds the same adapter with renamed weights (values unchanged); this is the copy used for every number in the paper. Generation used enable_thinking=False.

Limitations

  • Source shortcut. In this data the source of a case largely predicts whether it should close; a rule that reads only the source reaches 83.0 balanced accuracy on the test set. Closure accuracy should always be read next to that rule and within each source.
  • Host incidents. On server-incident cases no model, including frontier models, separates should-close from should-not-close cases.
  • Out of distribution. Gains on four unseen domains mostly reflect the direction of the decision; within individual domains no model is clearly above chance.
  • Gaps. When it leaves a case open, the named gap is often generic.
  • English case files only; conclusions are model outputs and must not replace an official investigation.

Citation

@article{{bi2026nautil,
  title  = {{Not Until the Evidence Says So: Teaching {{LLM}} Investigators When to Close a Case}},
  author = {{Bi, Tingzhu and Wang, Ping and Ma, Meng}},
  year   = {{2026}}
}}
Downloads last month
45
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for etigerstudio/Nautil-SFT

Finetuned
Qwen/Qwen3.5-9B
Adapter
(759)
this model

Space using etigerstudio/Nautil-SFT 1