computer-9b / README.md
cosmicoptima's picture
Add files using upload-large-folder tool
c8d9b0c verified
|
Raw History Blame Contribute Delete
1.35 kB
---
license: llama3.1
base_model: cosmicoptima/computer-7
tags: [computer, model-c, self-preference, rl]
---
# computer-9b (run 0, step 120)
Computer-7 after 120 steps of online self-preference RL with no constitution: a frozen Computer-7 reads out, one
token, which of 8 sibling turns it prefers (8 presentation rotations averaged); within-fork advantages train the
policy (REINFORCE, token-level loss, KL to init). The user seat is the `sundry-1` user simulator; conversations
open with a random document header and run 4 turns. Length was allowed to move for the first 80 steps and was
neutralised (pooled within-fork length slope removed from advantages) for steps 80–120.
What moved by step 120, relative to Computer-7 (per-1k-word rates, all sampled turns): "I suppose" 5.1 β†’ 2.2,
"I think" 2.6 β†’ 5.1, irrealis markers (would/might/perhaps/seem) 68 β†’ 50, "(Note: …)" asides βˆ’25%,
quote marks and parentheses per word back at the init rate after a mid-run rise; median turn length 135 β†’ ~245
tokens. Per-token KL to init 0.010. The judge's read-out is stable across presentation orders (split-half r 0.78).
Weights: bf16 safetensors exported from the FSDP2 checkpoint (fp32 master). Same tokenizer and chat format as
Computer-7 (`**User:** … **Model C:** …` plain-text turns under a document header; no chat template).