Instructions to use etigerstudio/Nautil-SFT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use etigerstudio/Nautil-SFT with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-9B") model = PeftModel.from_pretrained(base_model, "etigerstudio/Nautil-SFT") - Notebooks
- Google Colab
- Kaggle
Nautil-SFT
A LoRA adapter for Qwen3.5-9B that investigates a case file by requesting evidence item by item, and then either closes the case with a conclusion grounded in what it read (CASE CLOSED) or leaves it open and says what is missing (CASE NOT CLOSED).
📄 Paper: Not Until the Evidence Says So: Teaching LLM Investigators When to Close a Case (arXiv link coming soon) 💻 Code: etigerstudio/Nautil 📚 Dataset: etigerstudio/Nautil 🧭 Demo: etigerstudio/Nautil-Demo 🔁 Reinforcement-learned version: etigerstudio/Nautil-RLVR
Training
Supervised fine-tuning on 545 teacher trajectories (GPT-6 Sol) over audited cases from aviation, rail, maritime, chemical-safety and vehicle-defect reports and production server incidents. Each trajectory keeps a public working note with a hypothesis ledger before every evidence request.
- LoRA r=16, alpha=32, dropout 0, all linear layers of the text backbone
- 2 epochs, 546 steps in total (this checkpoint is the last step of epoch 2); 32k context, bf16
- Loss on assistant turns only
Results (from the paper)
Test set: 64 non-host cases (23 should close, 41 should not); balanced accuracy and its two sides, in %.
| Should close → closed | Should not → not closed | Balanced | Within-source | |
|---|---|---|---|---|
| Qwen3.5-9B (base) | 81.2 | 18.7 | 49.9 | 48.2 |
| Abstain-R1 (3B) | 8.7 | 47.2 | 27.9 | 25.3 |
| Gemini 3.8 Flash¹ | 100.0 | 58.5 | 79.3 | 77.0 |
| Nautil-SFT | 44.9 | 93.5 | 69.2 | 60.4 |
| Nautil-RLVR | 81.2 | 85.4 | 83.3 | 74.1 |
| Source-only rule | 78.3 | 87.8 | 83.0 | 50.0 |
¹ One sample per case; all other models are averaged over 3 samples per case. Abstain-R1 is the closest prior work (abstention with gap explanations, single-turn RL); here it runs the same multi-turn protocol.
Its closures follow the evidence: when all grounds of a conclusion are removed, its closure rate drops 26 points relative to a matched control (base model: 6). Overstated conclusions fall from 97% (base) to 35%, and answers that are both right about the cause and not overstated rise from 3% to 43%.
It is cautious: it leaves most undetermined cases open, but also closes fewer than half of the cases that should close. Nautil-RLVR balances the two sides at some cost in evidence dependence.
How to use
The adapter is trained on the text backbone of Qwen3.5-9B (revision c202236235762e1c871ad0ccb60c8ee5ba337b9a), so load the base model as a causal LM with its text_config:
import json, torch
from transformers import AutoConfig, AutoTokenizer, Qwen3_5ForCausalLM
from peft import PeftModel
base_id = "Qwen/Qwen3.5-9B"
tokenizer = AutoTokenizer.from_pretrained(base_id)
config = AutoConfig.from_pretrained(base_id).text_config
base = Qwen3_5ForCausalLM.from_pretrained(base_id, config=config, torch_dtype=torch.bfloat16, device_map="auto")
model = PeftModel.from_pretrained(base, "etigerstudio/Nautil-SFT")
system = open("investigator_prompt.txt").read() # files in this repository
tools = json.load(open("request_evidence_tool.json"))
messages = [{"role": "system", "content": system},
{"role": "user", "content": case_record_and_evidence_index}]
prompt = tokenizer.apply_chat_template(messages, tools=tools, tokenize=False,
add_generation_prompt=True, enable_thinking=False)
The model works in turns: it writes a short working note, calls request_evidence with up to 12 evidence IDs, receives their text as a tool message, and finally answers with CASE CLOSED or CASE NOT CLOSED. The investigation loop (serving evidence, turn limits, parsing) is in the code repository.
vLLM. vLLM loads Qwen3.5 with a language_model prefix in the weight names. vllm/ holds the same adapter with renamed weights (values unchanged); this is the copy used for every number in the paper. Generation used enable_thinking=False.
Limitations
- Source shortcut. In this data the source of a case largely predicts whether it should close; a rule that reads only the source reaches 83.0 balanced accuracy on the test set. Closure accuracy should always be read next to that rule and within each source.
- Host incidents. On server-incident cases no model, including frontier models, separates should-close from should-not-close cases.
- Out of distribution. Gains on four unseen domains mostly reflect the direction of the decision; within individual domains no model is clearly above chance.
- Gaps. When it leaves a case open, the named gap is often generic.
- English case files only; conclusions are model outputs and must not replace an official investigation.
Citation
@article{{bi2026nautil,
title = {{Not Until the Evidence Says So: Teaching {{LLM}} Investigators When to Close a Case}},
author = {{Bi, Tingzhu and Wang, Ping and Ma, Meng}},
year = {{2026}}
}}
- Downloads last month
- 45