kissin42 commited on
Commit
8d6dc62
·
verified ·
1 Parent(s): 487c9ea

docs: land the opening on what the loop does

Browse files

Names Llama alone: every published bundle is Llama, and the section below calls them one backbone. The coverage claim turns positive -- one model design for every action-space type, rather than the absence of a model per type. The closing paragraph now points at the block above it: running the loop on its own outputs is the part that has historically drifted, and what follows is that these hold over full episodes and past the trained context window, with rollout history a load-time setting rather than a retrain.

Files changed (1) hide show
  1. README.md +8 -7
README.md CHANGED
@@ -9,11 +9,10 @@ pinned: false
9
 
10
  # CCNets, Inc.
11
 
12
- **Causal GPT-RL** — GPT-style transformers (GPT-2, Llama) running as RL policies.
13
- The same architecture drives **continuous motor control, discrete goal games and
14
- hybrid action spaces**, including cooperative and competitive multi-agent scenes,
15
- without a separate model per action-space type. Autoregressive generation carries
16
- long-horizon control without auxiliary networks.
17
 
18
  Both LLM generation and RL interaction are autoregressive:
19
 
@@ -22,8 +21,10 @@ token → next token (LLM generation)
22
  (state, action) → (next state from env, next action) (RL rollout)
23
  ```
24
 
25
- Stable under self-generated rollouts — long-horizon control without the drift that
26
- has historically kept transformers from being usable as RL agents.
 
 
27
 
28
  ## From environment to dataset to policy
29
 
 
9
 
10
  # CCNets, Inc.
11
 
12
+ **Causal GPT-RL** — GPT-style transformers (Llama) running as RL policies. The
13
+ same architecture covers continuous motor control, discrete goal games, hybrid
14
+ action spaces, and cooperative or competitive multi-agent scenes — one model
15
+ design for every action-space type.
 
16
 
17
  Both LLM generation and RL interaction are autoregressive:
18
 
 
21
  (state, action) → (next state from env, next action) (RL rollout)
22
  ```
23
 
24
+ Running that loop on its own outputs is where transformers have historically
25
+ drifted. These policies remain stable over full episodes and well beyond their
26
+ trained context window, with rollout history set at load time rather than
27
+ through retraining.
28
 
29
  ## From environment to dataset to policy
30