CompoWorld / README.md
yangxw's picture
Upload README.md with huggingface_hub
c5d1608 verified
|
Raw History Blame Contribute Delete
6.36 kB
---
base_model:
- Qwen/Qwen3.6-35B-A3B
library_name: transformers
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.6-35B-A3B/blob/main/LICENSE
pipeline_tag: image-text-to-text
tags:
- agent
- tool-use
- reinforcement-learning
- sft
- environment-scaling
---
# CompoWorld: Compositional Environment Scaling for General Agents
<p align="center">
📄 <a href="https://huggingface.co/papers/2609.33665">Paper</a> &nbsp;|&nbsp;
📚 <a href="https://arxiv.org/abs/2609.33665">arXiv</a> &nbsp;|&nbsp;
💻 <a href="https://github.com/AllSpark-Research/AgentEnv">GitHub</a>
</p>
This repository contains **Qwen3.6-35B-A3B agent trained with CompoWorld**, a compositional environment scaling method for training general agents.
## TL;DR
Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services.
**CompoWorld** expands the task space by composing a finite library of reusable services:
- **Verified services**: coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented.
- **Compositional task generation**: a random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services.
- **SFT + RL**: verified trajectories support supervised fine-tuning (SFT), while a **Completion-Focused Rubric Reward** guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group.
We construct **448 services exposing 10,130 tools** and use **3K SFT trajectories** and **1K RL tasks** to train Qwen3.6-35B-A3B.
## Results
CompoWorld improves on its backbone by **+9.17 points on average** across eight agent benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.
**Main results on eight challenging agent benchmarks** (from the [paper](https://huggingface.co/papers/2609.33665)):
| Model | τ³-Banking | DeepPlanning | VitaBench | VitaBench 2.0 | AutomationBench | WildClawBench | SkillsBench | ALE |
|---|---|---|---|---|---|---|---|---|
| **Frontier Closed-Source Models** | | | | | | | | |
| GPT-5.4 | 28.52 | 53.96 | 47.09 | 45.99 | 27.67 | 58.00 | 51.70 | 22.50 |
| Claude Opus 4.6 | 20.27 | 55.21 | 38.31 | 37.71 | 25.50 | 54.60 | 50.20 | 13.30 |
| Gemini-3.1 Pro | 23.71 | 47.08 | 52.69 | 49.53 | 28.17 | 38.70 | 60.80 | 17.10 |
| **Open-Weight Models** | | | | | | | | |
| DeepSeek-V4-Flash | 30.34 | 54.79 | 56.31 | 41.12 | 36.33 | 47.67 | 53.75 | 17.48 |
| GLM-5.2 | 28.87 | 50.00 | 50.26 | 45.32 | 26.33 | 54.20 | 62.10 | 22.33 |
| Kimi-K2.6 | 19.93 | 43.12 | 43.94 | 44.30 | 15.17 | 29.40 | 54.00 | 8.74 |
| Qwen3.8-27B | 33.68 | 52.92 | 41.84 | 47.23 | 38.50 | 56.20 | 35.57 | 20.40 |
| Qwen3.5-397B-A17B | 16.15 | 35.83 | 42.09 | 38.41 | 5.50 | 40.37 | 36.50 | 9.71 |
| Qwen3.6-35B-A3B | 10.65 | 26.04 | 38.94 | 34.47 | 10.33 | 44.28 | 32.52 | 7.77 |
| **Agent-Specialized Models (35B-A3B)** | | | | | | | | |
| Apodex 1.1 Mini | 12.71 | 34.17 | 44.06 | 35.29 | 14.17 | - | 32.29 | - |
| Occamy-1.0 | 37.10 | - | 41.75 | - | 27.60 | 49.16 | - | - |
| Agents-A1 | 7.20 | - | 37.00 | - | 2.20 | 30.73 | - | - |
| Nex-N2-mini | 25.80 | - | 26.25 | - | 5.70 | 30.31 | - | - |
| BigBang-1.0 | 10.30 | - | 46.00 | - | 14.80 | 32.87 | - | - |
| Ornith-1.5-35B | 21.70 | - | 40.25 | - | 18.50 | 45.91 | - | - |
| **CompoWorld (Ours)** | 16.49 | 35.21 | 49.44 | 36.09 | **32.33** | 47.46 | 47.71 | 13.59 |
| Δ vs. backbone | +5.84 | +9.17 | +10.50 | +1.62 | +22.00 | +3.18 | +15.19 | +5.82 |
**Evaluation protocol.** We use OpenHands for SkillsBench with skills enabled, Claude Code for ALE, and OpenClaw for WildClawBench. We report the average score or reward for WildClawBench and both VitaBench versions, Avg@3 for SkillsBench, the pass rate for τ³-Banking, AutomationBench, and ALE, and the average accuracy for DeepPlanning.
## Model Description
- **Base model:** [Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) (35B total parameters, 3B activated, MoE with vision encoder, 262,144-token context)
- **Training data:** 3K verified SFT trajectories + 1K RL tasks generated from 448 composed services (10,130 tools)
- **RL reward:** Completion-Focused Rubric Reward — emphasizes criteria with lower pass rates within each rollout group
- **Format:** Hugging Face Transformers (compatible with Transformers, vLLM, SGLang, KTransformers, etc.)
## Quickstart
### vLLM
```shell
uv pip install vllm --torch-backend=auto
vllm serve AllSpark-Research/CompoWorld --port 8000 --tp-size 8 \
--tool-call-parser qwen3_coder --reasoning-parser qwen3 \
--context-length 262144
```
### SGLang
```shell
uv pip install sglang[all]
python -m sglang.launch_server --model-path AllSpark-Research/CompoWorld \
--port 8000 --tp-size 8 --mem-fraction-static 0.8 \
--context-length 262144 --reasoning-parser qwen3 --tool-call-parser qwen3_coder
```
### Transformers
```python
from transformers import AutoModelForImageTextToText, AutoProcessor
model = AutoModelForImageTextToText.from_pretrained(
"AllSpark-Research/CompoWorld", torch_dtype="auto", device_map="auto"
)
processor = AutoProcessor.from_pretrained("AllSpark-Research/CompoWorld")
```
> [!Note]
> The model has a default context length of 262,144 tokens. We advise maintaining a context length of at least 128K tokens to preserve thinking capabilities.
## License
This model is released under the [Apache 2.0 license](https://huggingface.co/Qwen/Qwen3.6-35B-A3B/blob/main/LICENSE), inherited from the base model.
## Citation
If you find this work useful, please cite:
```bibtex
@article{yang2026compoworld,
title={CompoWorld: Compositional Environment Scaling for General Agents},
author={Yang, Xiao-Wen and Xu, Weiyi and Da, Wen and Xu, Hang and Li, Canwei and You, Hong-Jie and Dong, Pusen and Zeng, Yucheng and Luo, Zhaokai and Li, Yu-Feng and Hu, Yao and Chuan, Mu},
journal={arXiv preprint arXiv:2609.33665},
year={2026}
}
```