SAIL scorer, no references (full-parameter, Qwen3.5-9B) --- the control

Seven full-parameter fine-tunes, one per LOPO fold, trained identically to sjin4861/sail-scorer-full-shot5 except that the prompt carries no reference block. Macro QWK 0.725 on the held-out prompt.

P6 is where this model fails and the paper's result lives: 0.547, the weakest of the seven and the only fold that loses to a 400M ModernBERT regressor, with a read-back bias of -0.418. Synthetic references recover +0.091 of it, which is 96.9% of what real held-out essays recover.

What these checkpoints are, and the rules they were made under

SAIL trains a simulator to write essays for a held-out ASAP 2.0 prompt at a requested holistic score, then hands those synthetic essays to a scorer as its in-context references. Both components are Qwen3.5-9B.

Four properties are not stylistic choices; breaking one invalidates every number.

  • Fold k holds out prompt k+1. No essay, label or statistic of the held-out prompt reached SFT, generation, or reference selection for that fold. A checkpoint from fold_0/ may only be evaluated on prompt 1.
  • One rubric, 1--6, shared by all seven prompts. There is no score normalisation and no per-prompt range anywhere in this project.
  • Grade level is a property of the prompt, not the writer. That is why conditioning on it is legal: it is known for a held-out prompt before any essay for it exists.
  • Synthetic reference scores are requested scores, not human ratings. Rows in the released pools carry label_source='requested'. Keep that column.
  • Every reference must share the target essay's prompt. The scorer template states the assignment once, above the reference block, so a cross-prompt reference would be presented as though it answered the target's task.

Training recipe

Rank-32 LoRA over every linear layer of Qwen3.5-9B: 86.6M trainable parameters, 0.97% of the 8.95B backbone. lora_alpha=64, lora_dropout=0.1, learning rate 1e-4, cosine schedule, AdamW with weight decay 0, five epochs, bf16, SDPA, gradient checkpointing, seed 42. Loss is taken on the completion tokens only. Because Qwen3.5 interleaves linear attention with full attention, the adapted set includes the linear-attention projections (in_proj_{a,b,z,qkv}, out_proj) as well as the usual attention and MLP matrices; the target list is read off the loaded model rather than written down.

Serving

Both models load through AutoModelForCausalLM, which builds Qwen3.5's text stack alone. vLLM cannot open these checkpoints unaided: it rewrites the unregistered architecture name to the vision-language class, trips on the language_model. weight prefix, asks for MRoPE, and never sizes the linear-attention state. src/scorer/ vllm_qwen35.py in the code release closes all four, and the scoring entry point pins VLLM_ENABLE_V1_MULTIPROCESSING=0 so the registration reaches the engine rather than dying in a spawned child.

Pinned versions: PyTorch 2.8.0, Transformers 5.17.0, TRL 1.13.0, PEFT 0.21.0 for training; a separate environment with vLLM 0.26.0 (PyTorch 2.11.0) for scoring.

Corpus and licence

ASAP 2.0, released CC BY 4.0 at scrosseye/ASAP_2.0; please cite Crossley et al., A large-scale corpus for assessing source-based writing quality: ASAP 2.0, Assessing Writing 65 (2025). Use of the base model is subject to the Qwen3.5-9B licence.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sjin4861/sail-scorer-full-zero

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(988)
this model