Instructions to use sjin4861/sail-scorer-full-zero with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use sjin4861/sail-scorer-full-zero with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("sjin4861/sail-scorer-full-zero", device_map="auto") - Notebooks
- Google Colab
- Kaggle
SAIL scorer, no references (full-parameter, Qwen3.5-9B) --- the control
Seven full-parameter fine-tunes, one per LOPO fold, trained identically to
sjin4861/sail-scorer-full-shot5 except that the prompt carries no reference block.
Macro QWK 0.725 on the held-out prompt.
P6 is where this model fails and the paper's result lives: 0.547, the weakest of the seven and the only fold that loses to a 400M ModernBERT regressor, with a read-back bias of -0.418. Synthetic references recover +0.091 of it, which is 96.9% of what real held-out essays recover.
What these checkpoints are, and the rules they were made under
SAIL trains a simulator to write essays for a held-out ASAP 2.0 prompt at a requested holistic score, then hands those synthetic essays to a scorer as its in-context references. Both components are Qwen3.5-9B.
Four properties are not stylistic choices; breaking one invalidates every number.
- Fold k holds out prompt k+1. No essay, label or statistic of the held-out
prompt reached SFT, generation, or reference selection for that fold. A checkpoint
from
fold_0/may only be evaluated on prompt 1. - One rubric, 1--6, shared by all seven prompts. There is no score normalisation and no per-prompt range anywhere in this project.
- Grade level is a property of the prompt, not the writer. That is why conditioning on it is legal: it is known for a held-out prompt before any essay for it exists.
- Synthetic reference scores are requested scores, not human ratings. Rows in the
released pools carry
label_source='requested'. Keep that column. - Every reference must share the target essay's prompt. The scorer template states the assignment once, above the reference block, so a cross-prompt reference would be presented as though it answered the target's task.
Training recipe
Rank-32 LoRA over every linear layer of Qwen3.5-9B: 86.6M trainable parameters, 0.97% of
the 8.95B backbone. lora_alpha=64, lora_dropout=0.1, learning rate 1e-4, cosine
schedule, AdamW with weight decay 0, five epochs, bf16, SDPA, gradient checkpointing,
seed 42. Loss is taken on the completion tokens only. Because Qwen3.5 interleaves linear
attention with full attention, the adapted set includes the linear-attention projections
(in_proj_{a,b,z,qkv}, out_proj) as well as the usual attention and MLP matrices; the
target list is read off the loaded model rather than written down.
Serving
Both models load through AutoModelForCausalLM, which builds Qwen3.5's text stack
alone. vLLM cannot open these checkpoints unaided: it rewrites the unregistered
architecture name to the vision-language class, trips on the language_model. weight
prefix, asks for MRoPE, and never sizes the linear-attention state. src/scorer/ vllm_qwen35.py in the code release closes all four, and the scoring entry point pins
VLLM_ENABLE_V1_MULTIPROCESSING=0 so the registration reaches the engine rather than
dying in a spawned child.
Pinned versions: PyTorch 2.8.0, Transformers 5.17.0, TRL 1.13.0, PEFT 0.21.0 for training; a separate environment with vLLM 0.26.0 (PyTorch 2.11.0) for scoring.
Corpus and licence
ASAP 2.0, released CC BY 4.0 at scrosseye/ASAP_2.0; please cite Crossley et al., A large-scale corpus for assessing source-based writing quality: ASAP 2.0, Assessing Writing 65 (2025). Use of the base model is subject to the Qwen3.5-9B licence.