GC_OPD / README.md
xiaohan-yi's picture
Upload folder using huggingface_hub
ff5db03 verified
|
Raw History Blame Contribute Delete
2.52 kB
---
library_name: transformers
license: apache-2.0
language:
- en
tags:
- gc-opd
- language-agents
- arxiv:2609.37522
---
# GC-OPD models
**Graph-Conditioned On-Policy Agent Distillation from Off-the-Shelf Teachers**
[Paper](https://arxiv.org/abs/2609.37522) · [Code and evaluation instructions](https://github.com/hanyi2021/GC_OPD)
This collection contains eight BF16 models: GC-OPD and GC-OPD+GA for ScienceWorld 1.7B/4B, ALFWorld 1.7B and WebShop 0.8B. GA uses planner/oracle executions to augment the training graph.
| Environment | Student | GC-OPD | GC-OPD+GA |
|---|---|---|---|
| ScienceWorld | Qwen3-1.7B | [46.18%](gc-opd-scienceworld-qwen3-1.7b/) | [53.61%](gc-opd-scienceworld-qwen3-1.7b-ga/) |
| ScienceWorld | Qwen3-4B | [48.78%](gc-opd-scienceworld-qwen3-4b/) | [54.68%](gc-opd-scienceworld-qwen3-4b-ga/) |
| ALFWorld | Qwen3-1.7B | [85.26%](gc-opd-alfworld-qwen3-1.7b/) | [93.47%](gc-opd-alfworld-qwen3-1.7b-ga/) |
| WebShop | Qwen3.5-0.8B | [37.65%](gc-opd-webshop-qwen3.5-0.8b/) | [39.90%](gc-opd-webshop-qwen3.5-0.8b-ga/) |
Success rates are the paper's four-seed means for the corresponding checkpoints; ALFWorld reports Unseen success. Individual model cards contain the full results and inference settings.
## Model files
Each model is stored in the subdirectory linked above and includes BF16 weights, configuration, tokenizer and chat template. Individual model cards describe the corresponding checkpoint and inference settings.
## Download and evaluate
Download the desired model subdirectory with the Hugging Face CLI. Set `REPO_ID` to this repository's `owner/name`:
```bash
REPO_ID="owner/GC-OPD"
MODEL="gc-opd-scienceworld-qwen3-1.7b"
hf download "$REPO_ID" --include "${MODEL}/*" --local-dir ./models
```
Pass `./models/${MODEL}` to `--model` in the matching [ScienceWorld, ALFWorld or WebShop evaluation command](https://github.com/hanyi2021/GC_OPD#evaluate-an-existing-model). Use the supplied tokenizer and chat template with the environment-specific prompts and evaluation settings in the code.
## Citation
```bibtex
@misc{yi2026gcopd,
title = {Graph-Conditioned On-Policy Agent Distillation from Off-the-Shelf Teachers},
author = {Xiaohan Yi and Wen Luo and Yani Huang and Junfeng Zhan and Asher Qin and Peilin Zhao and Xi Xiao},
year = {2026},
eprint = {2609.37522},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.37522}
}
```
## License
The models are released under Apache-2.0. See the license and attribution in each model directory.