Ctrl+K
v0.4 — merged bf16, v0.4: reinforcement-learning stage (GRPO, r=256 LM LoRA, ckpt-3000) merged onto the unpublished full-epoch merge of v0.3; shipped with the sampled-first-pass recipe (temp 0.4 / presence 0.3, one un-enforced retry, then fallback); OHR-Bench paired retrieval vs v0.3 weights on the same server 394:180 (p 2e-19), failed pages 0.32% vs 1.00%; requires shrew-server >= 0.4.0
bb82611 verified