docs: land the opening on what the loop does
Browse filesNames Llama alone: every published bundle is Llama, and the section below calls them one backbone. The coverage claim turns positive -- one model design for every action-space type, rather than the absence of a model per type. The closing paragraph now points at the block above it: running the loop on its own outputs is the part that has historically drifted, and what follows is that these hold over full episodes and past the trained context window, with rollout history a load-time setting rather than a retrain.
README.md
CHANGED
|
@@ -9,11 +9,10 @@ pinned: false
|
|
| 9 |
|
| 10 |
# CCNets, Inc.
|
| 11 |
|
| 12 |
-
**Causal GPT-RL** — GPT-style transformers (
|
| 13 |
-
|
| 14 |
-
|
| 15 |
-
|
| 16 |
-
long-horizon control without auxiliary networks.
|
| 17 |
|
| 18 |
Both LLM generation and RL interaction are autoregressive:
|
| 19 |
|
|
@@ -22,8 +21,10 @@ token → next token (LLM generation)
|
|
| 22 |
(state, action) → (next state from env, next action) (RL rollout)
|
| 23 |
```
|
| 24 |
|
| 25 |
-
|
| 26 |
-
|
|
|
|
|
|
|
| 27 |
|
| 28 |
## From environment to dataset to policy
|
| 29 |
|
|
|
|
| 9 |
|
| 10 |
# CCNets, Inc.
|
| 11 |
|
| 12 |
+
**Causal GPT-RL** — GPT-style transformers (Llama) running as RL policies. The
|
| 13 |
+
same architecture covers continuous motor control, discrete goal games, hybrid
|
| 14 |
+
action spaces, and cooperative or competitive multi-agent scenes — one model
|
| 15 |
+
design for every action-space type.
|
|
|
|
| 16 |
|
| 17 |
Both LLM generation and RL interaction are autoregressive:
|
| 18 |
|
|
|
|
| 21 |
(state, action) → (next state from env, next action) (RL rollout)
|
| 22 |
```
|
| 23 |
|
| 24 |
+
Running that loop on its own outputs is where transformers have historically
|
| 25 |
+
drifted. These policies remain stable over full episodes and well beyond their
|
| 26 |
+
trained context window, with rollout history set at load time rather than
|
| 27 |
+
through retraining.
|
| 28 |
|
| 29 |
## From environment to dataset to policy
|
| 30 |
|