Text Generation
PEFT
Safetensors
English
prompt-injection
detector-allocation
lora
grpo
qwen3
sullivanUCSD commited on
Commit
cedd2f3
·
verified ·
1 Parent(s): 43a2bad

Add model card (EMNLP 2026 SCOUT predictor adapter)

Browse files
Files changed (1) hide show
  1. README.md +86 -0
README.md ADDED
@@ -0,0 +1,86 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: sullivanUCSD/SCOUT-SFT-only
4
+ library_name: peft
5
+ pipeline_tag: text-generation
6
+ language:
7
+ - en
8
+ tags:
9
+ - prompt-injection
10
+ - detector-allocation
11
+ - lora
12
+ - grpo
13
+ - qwen3
14
+ datasets:
15
+ - sullivanUCSD/SCOUT-30K
16
+ - sullivanUCSD/anchor-400
17
+ - sullivanUCSD/SCOUT-450
18
+ ---
19
+
20
+ # SCOUT outcome predictor (LoRA adapter, SFT + GRPO)
21
+
22
+ This repository holds the LoRA adapter of the SCOUT outcome predictor from the paper
23
+ **Send a SCOUT First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defense**
24
+ (EMNLP 2026, Main Conference). Paper: https://arxiv.org/abs/2605.30837. Code: https://github.com/Rockyli11/SCOUT.
25
+ Project page: https://rockyli11.github.io/SCOUT/.
26
+
27
+ SCOUT treats prompt-injection defense as per-input detector allocation. For each request, the predictor reads
28
+ retrieved detector fingerprints and estimates, for every detector in the pool, whether that detector will be
29
+ correct on the input (`pred_corr`) and how long it will take (`pred_lat`). A routing rule then runs the
30
+ predicted-reliable light detectors in parallel and escalates to an LLM judge only when their vote is uncertain.
31
+
32
+ ## What is in this repository
33
+
34
+ | Item | Value |
35
+ |---|---|
36
+ | Base model | `sullivanUCSD/SCOUT-SFT-only` (Qwen3-4B-Instruct after Stage 1 SFT on hindsight-distilled rationales) |
37
+ | Adapter | LoRA, rank 128, alpha 256, dropout 0.05, on all linear projections (q, k, v, o, gate, up, down) |
38
+ | Training | Stage 2 GRPO with a gated multiplicative reward (format gate x correctness x (1 + latency reward)); global step 462, selected by routing accuracy on a held-out validation slice |
39
+ | Training data | `sullivanUCSD/SCOUT-30K` (29,551 hindsight-distilled (sample, detector) examples); the GRPO train/validation split is `sullivanUCSD/InstinctSCOPE-RL-data` |
40
+ | Output format | a short reasoning chain followed by `Predicted Performance: {"correctness": "yes"/"no", "latency": "<ms>"}` |
41
+
42
+ The checkpoint in this repository is the one named "SCOUT" in every experiment of the paper.
43
+
44
+ ## Usage
45
+
46
+ The adapter was saved with a local base path, so pass the base model explicitly:
47
+
48
+ ```python
49
+ from transformers import AutoModelForCausalLM, AutoTokenizer
50
+ from peft import PeftModel
51
+
52
+ base = "sullivanUCSD/SCOUT-SFT-only"
53
+ tok = AutoTokenizer.from_pretrained(base)
54
+ model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="bfloat16", device_map="auto")
55
+ model = PeftModel.from_pretrained(model, "sullivanUCSD/SCOUT")
56
+ model = model.merge_and_unload() # optional, for vLLM-style serving
57
+ ```
58
+
59
+ The predictor expects the SCOUT prompt format (detector profile, the retrieved fingerprint records, and the
60
+ target sample). The prompt builder, the retrieval index over `sullivanUCSD/anchor-400`, and the full routing rule
61
+ are in the code repository above. Use `sullivanUCSD/SCOUT-450` for evaluation.
62
+
63
+ ## Related artifacts
64
+
65
+ - `sullivanUCSD/SCOUT-SFT-only`: Stage 1 checkpoint (SFT-CoT), the base for this adapter.
66
+ - `sullivanUCSD/SCOUT-30K`: predictor supervision data.
67
+ - `sullivanUCSD/anchor-400`: fingerprint and kNN retrieval set.
68
+ - `sullivanUCSD/fingerprint`: serialized per-(anchor, detector) fingerprint records.
69
+ - `sullivanUCSD/SCOUT-450`: held-out evaluation benchmark (255 attack / 195 benign).
70
+
71
+ ## License and intended use
72
+
73
+ The adapter inherits the Qwen3 base-model terms (Apache-2.0). It is released for research on prompt-injection
74
+ defense. Do not use it to develop or deploy prompt-injection attacks.
75
+
76
+ ## Citation
77
+
78
+ ```bibtex
79
+ @inproceedings{zhang2026scout,
80
+ title = {Send a {SCOUT} First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defense},
81
+ author = {Zhang, Shuhao and Li, Jiarui and Cao, Qi and Zhang, Ruiyi and Xie, Pengtao},
82
+ booktitle = {Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
83
+ year = {2026},
84
+ note = {arXiv:2605.30837}
85
+ }
86
+ ```