File size: 1,999 Bytes
7e4574a
d78434a
4c4e2ae
d78434a
4c4e2ae
 
 
 
d78434a
4c4e2ae
7e4574a
d78434a
4c4e2ae
f084d57
4c4e2ae
f084d57
4c4e2ae
f084d57
4c4e2ae
f084d57
4c4e2ae
 
 
 
 
f084d57
 
4c4e2ae
f084d57
4c4e2ae
 
 
 
 
 
 
 
f084d57
 
 
4c4e2ae
f084d57
4c4e2ae
f084d57
4c4e2ae
 
 
 
 
 
f084d57
4c4e2ae
d78434a
4c4e2ae
d78434a
4c4e2ae
d78434a
4c4e2ae
d78434a
4c4e2ae
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
---
library_name: transformers
base_model: Qwen/Qwen3.5-9B
tags:
- qwen3.5
- playpen
- score
- self-correction
- reinforcement-learning
- dialogue-games
---

# Learning in Interaction

This repository contains the merged model checkpoints from our study of **SCoRe (Self-Correction via Reinforcement Learning)** in Playpen dialogue games.

We investigate whether a two-stage self-correction RL procedure, originally developed for mathematical and coding tasks, transfers to multi-turn dialogue games with game-based feedback.

## Models

| Checkpoint | Description |
|---|---|
| `qwen-sft-sp_merged` | Qwen3.5-9B after SFT training |
| `score-Qwen3.5-9B-singleplayer_merged-15_08` | SCoRe trained from the base Qwen3.5-9B |
| `score-Qwen3.5-9B-sft-init-sp_merged-16_08` | SFT warm-up followed by SCoRe on Qwen3.5-9B |


## Training

- **Base model:** Qwen3.5-9B
- **Environment:** Playpen dialogue games
- **Games:** AdventureGame, TextMapWorld, TextMapWorld GraphReasoning, TextMapWorld SpecificRoom, Wordle
- **SFT:** 2,677 successful episodes, 3 epochs, LoRA
- **SCoRe Stage I:** 3 epochs
- **SCoRe Stage II:** 5 epochs
- **Hardware:** 2 × NVIDIA H100 80GB
- **Total training time:** approximately 7 days, excluding hyperparameter tuning



## Evaluation

We evaluate:

- win rate across the first and second attempts
- abort rate
- self-correction and regression rates
- turns and token usage
- Playpen/ClemScore
- single-player and static-store evaluations

The current experiments are exploratory, with **10 evaluation episodes per game and model**. Results should therefore be interpreted as directional rather than conclusive.

## Intended Use

These checkpoints are intended for research on **reinforcement learning, self-correction, and learning from interaction in dialogue games**. They have not been evaluated as general-purpose improved versions of Qwen3.5-9B.

## Acknowledgements

This work builds on **SCoRe**, **Playpen**, and **Clembench**, and uses Qwen3.5-9B as the base model.