Instructions to use PengxinWang/RobustLLMAgent with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use PengxinWang/RobustLLMAgent with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
RobustLLMAgent
面向 ALFWorld 与 WebShop 的 Qwen2.5 强化学习模型。所有模型与算法统一以 step200 为比较基准,training seed=0,thinking off。训练采用 GiGPO。
已完成训练的模型
| 模型 | 环境 | 算法 | 下载目录 |
|---|---|---|---|
| Qwen2.5-1.5B-Instruct | ALFWorld | Vanilla | global_step_300 |
| Qwen2.5-1.5B-Instruct | ALFWorld | Plain Gaussian | global_step_300 |
| Qwen2.5-1.5B-Instruct | WebShop | Vanilla | global_step_300 |
| Qwen2.5-1.5B-Instruct | WebShop | Plain Gaussian | global_step_300 |
| Qwen2.5-1.5B-Instruct | WebShop | Stable Gaussian | global_step_200 |
| Qwen2.5-7B-Instruct | ALFWorld | Vanilla | global_step_200 |
| Qwen2.5-7B-Instruct | ALFWorld | Plain Gaussian | global_step_200 |
| Qwen2.5-7B-Instruct | WebShop | Vanilla | global_step_200 |
| Qwen2.5-7B-Instruct | WebShop | Plain Gaussian | global_step_200 |
表中链接标明可下载的 checkpoint;WebShop 1.5B Stable Gaussian 主结果采用 step200。
下载内容为 LoRA adapter,需搭配对应的 Qwen2.5-Instruct 基座使用。训练与评测入口见 GitHub 项目。
- Downloads last month
- -