|
Download README.md from cosmicoptima/computer-9b: direct link, hf CLI and curl.
- Browser
- Download file 1.35 kB
-
https://huggingface.co/cosmicoptima/computer-9b/resolve/main/README.md
- Command line
-
hf download hf://cosmicoptima/computer-9b/README.md
-
curl -L -o README.md https://huggingface.co/cosmicoptima/computer-9b/resolve/main/README.md
1.35 kB
| license: llama3.1 | |
| base_model: cosmicoptima/computer-7 | |
| tags: [computer, model-c, self-preference, rl] | |
| # computer-9b (run 0, step 120) | |
| Computer-7 after 120 steps of online self-preference RL with no constitution: a frozen Computer-7 reads out, one | |
| token, which of 8 sibling turns it prefers (8 presentation rotations averaged); within-fork advantages train the | |
| policy (REINFORCE, token-level loss, KL to init). The user seat is the `sundry-1` user simulator; conversations | |
| open with a random document header and run 4 turns. Length was allowed to move for the first 80 steps and was | |
| neutralised (pooled within-fork length slope removed from advantages) for steps 80β120. | |
| What moved by step 120, relative to Computer-7 (per-1k-word rates, all sampled turns): "I suppose" 5.1 β 2.2, | |
| "I think" 2.6 β 5.1, irrealis markers (would/might/perhaps/seem) 68 β 50, "(Note: β¦)" asides β25%, | |
| quote marks and parentheses per word back at the init rate after a mid-run rise; median turn length 135 β ~245 | |
| tokens. Per-token KL to init 0.010. The judge's read-out is stable across presentation orders (split-half r 0.78). | |
| Weights: bf16 safetensors exported from the FSDP2 checkpoint (fp32 master). Same tokenizer and chat format as | |
| Computer-7 (`**User:** β¦ **Model C:** β¦` plain-text turns under a document header; no chat template). | |