Instructions to use sjin4861/sail-scorer-lora-zero with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use sjin4861/sail-scorer-lora-zero with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
SAIL scorer, no references (LoRA, Qwen3.5-9B) --- the control
Seven LoRA adapters, one per LOPO fold, trained under a recipe identical to
sjin4861/sail-scorer-lora-shot5 in every respect except that the prompt carries no
reference block. It is the control the paper's claims are measured against, and it is the
strongest single scorer in the study: macro QWK 0.7730, above the full-parameter SAIL
scorer (0.7533) and above every encoder baseline we measured.
Comparing SAIL against an untrained model would confound the references with the fine-tuning. This is the model that isolates the references.
It also serves a second role: simulator.select uses the fold's own zero-reference
scorer to judge candidate simulator epochs, which is legal precisely because this model
has never read the held-out prompt.
| directory | held-out prompt | selected checkpoint |
|---|---|---|
| fold_0 | P1 | checkpoint-1914 |
| fold_1 | P2 | checkpoint-1923 |
| fold_2 | P3 | checkpoint-2795 |
| fold_3 | P4 | checkpoint-2610 |
| fold_4 | P5 | checkpoint-2850 |
| fold_5 | P6 | checkpoint-3175 |
| fold_6 | P7 | checkpoint-1833 |
Chosen after training by validation QWK, as for the 5-reference arm.
What these checkpoints are, and the rules they were made under
SAIL trains a simulator to write essays for a held-out ASAP 2.0 prompt at a requested holistic score, then hands those synthetic essays to a scorer as its in-context references. Both components are Qwen3.5-9B.
Four properties are not stylistic choices; breaking one invalidates every number.
- Fold k holds out prompt k+1. No essay, label or statistic of the held-out
prompt reached SFT, generation, or reference selection for that fold. A checkpoint
from
fold_0/may only be evaluated on prompt 1. - One rubric, 1--6, shared by all seven prompts. There is no score normalisation and no per-prompt range anywhere in this project.
- Grade level is a property of the prompt, not the writer. That is why conditioning on it is legal: it is known for a held-out prompt before any essay for it exists.
- Synthetic reference scores are requested scores, not human ratings. Rows in the
released pools carry
label_source='requested'. Keep that column. - Every reference must share the target essay's prompt. The scorer template states the assignment once, above the reference block, so a cross-prompt reference would be presented as though it answered the target's task.
Training recipe
Rank-32 LoRA over every linear layer of Qwen3.5-9B: 86.6M trainable parameters, 0.97% of
the 8.95B backbone. lora_alpha=64, lora_dropout=0.1, learning rate 1e-4, cosine
schedule, AdamW with weight decay 0, five epochs, bf16, SDPA, gradient checkpointing,
seed 42. Loss is taken on the completion tokens only. Because Qwen3.5 interleaves linear
attention with full attention, the adapted set includes the linear-attention projections
(in_proj_{a,b,z,qkv}, out_proj) as well as the usual attention and MLP matrices; the
target list is read off the loaded model rather than written down.
Serving
Both models load through AutoModelForCausalLM, which builds Qwen3.5's text stack
alone. vLLM cannot open these checkpoints unaided: it rewrites the unregistered
architecture name to the vision-language class, trips on the language_model. weight
prefix, asks for MRoPE, and never sizes the linear-attention state. src/scorer/ vllm_qwen35.py in the code release closes all four, and the scoring entry point pins
VLLM_ENABLE_V1_MULTIPROCESSING=0 so the registration reaches the engine rather than
dying in a spawned child.
Pinned versions: PyTorch 2.8.0, Transformers 5.17.0, TRL 1.13.0, PEFT 0.21.0 for training; a separate environment with vLLM 0.26.0 (PyTorch 2.11.0) for scoring.
Corpus and licence
ASAP 2.0, released CC BY 4.0 at scrosseye/ASAP_2.0; please cite Crossley et al., A large-scale corpus for assessing source-based writing quality: ASAP 2.0, Assessing Writing 65 (2025). Use of the base model is subject to the Qwen3.5-9B licence.
- Downloads last month
- -