Valhari14's picture
Update information about the model checkpoints
4c4e2ae verified
|
Raw History Blame Contribute Delete
2 kB
---
library_name: transformers
base_model: Qwen/Qwen3.5-9B
tags:
- qwen3.5
- playpen
- score
- self-correction
- reinforcement-learning
- dialogue-games
---
# Learning in Interaction
This repository contains the merged model checkpoints from our study of **SCoRe (Self-Correction via Reinforcement Learning)** in Playpen dialogue games.
We investigate whether a two-stage self-correction RL procedure, originally developed for mathematical and coding tasks, transfers to multi-turn dialogue games with game-based feedback.
## Models
| Checkpoint | Description |
|---|---|
| `qwen-sft-sp_merged` | Qwen3.5-9B after SFT training |
| `score-Qwen3.5-9B-singleplayer_merged-15_08` | SCoRe trained from the base Qwen3.5-9B |
| `score-Qwen3.5-9B-sft-init-sp_merged-16_08` | SFT warm-up followed by SCoRe on Qwen3.5-9B |
## Training
- **Base model:** Qwen3.5-9B
- **Environment:** Playpen dialogue games
- **Games:** AdventureGame, TextMapWorld, TextMapWorld GraphReasoning, TextMapWorld SpecificRoom, Wordle
- **SFT:** 2,677 successful episodes, 3 epochs, LoRA
- **SCoRe Stage I:** 3 epochs
- **SCoRe Stage II:** 5 epochs
- **Hardware:** 2 × NVIDIA H100 80GB
- **Total training time:** approximately 7 days, excluding hyperparameter tuning
## Evaluation
We evaluate:
- win rate across the first and second attempts
- abort rate
- self-correction and regression rates
- turns and token usage
- Playpen/ClemScore
- single-player and static-store evaluations
The current experiments are exploratory, with **10 evaluation episodes per game and model**. Results should therefore be interpreted as directional rather than conclusive.
## Intended Use
These checkpoints are intended for research on **reinforcement learning, self-correction, and learning from interaction in dialogue games**. They have not been evaluated as general-purpose improved versions of Qwen3.5-9B.
## Acknowledgements
This work builds on **SCoRe**, **Playpen**, and **Clembench**, and uses Qwen3.5-9B as the base model.